Build one hard, real-world task in your field, used to test how well AI agents handle expert-level technical work.
Each task is open-ended: there is no single right answer, and an automatic scorer grades the AI agent's result from 0 to 1 with partial credit. For example: "Build the fastest possible attention engine."