Resources
Frequently Asked Questions
Common questions about tool and software access, task design, evaluation, authorship, venues, and timeline.
Authorship
Berkeley RDI is now expanding Agents' Last Exam toward a Nature-family submission. New contributors whose accepted work meets the bar below will be invited to join as co-authors on this submission.
Agents' Last Exam authorship is open to contributors across two tracks:
- Engineering track: execution & infrastructure support and new data pipeline build support. If you'd like to contribute here, contact sunyiyou@berkeley.edu & hanxinyang@berkeley.edu.
- Data track: providing task content and evaluation assets — no coding required (task specification, input files/assets, reference outputs/ground truth, objective rubric/QA pairs, domain expertise needed for verification).
All submitted tasks that pass review will be recognized and incorporated, with contributors listed under “Data Contributor” in the author list.
How we determine whether a task qualifies
Passing the automated AI agent review is highly recommended but not mandatory — the AI can occasionally make mistakes. That said, in our experience over 95% of tasks that pass the AI review also clear our internal evaluation. Every submission is checked against three control mechanisms:
- AI agent review— an automated first pass that gives a vibe- or topic-level read on whether your task is well-formed and in scope. It's a helpful signal rather than a hard gate, and we prioritize implementing tasks that score highly here.
- Quality check — we verify that your reference outputs are genuine, i.e., produced by an actual run of the workflow rather than fabricated. If we find a violation, we will withdraw the associated authorship.
- Difficulty & quantity check — each task is classified by difficulty, and authorship requires meeting either of the following thresholds:
- Near-term level: ≥ 3 tasks (additional variants of one task count as 1/3 of a task), or
- Last-exam level: ≥ 1 task.
After you submit
The automated agent review runs first, then our team reviews the submission for completeness and viability. You may be contacted for additional assets or clarifications. Task collection is ongoing, and the author list on the arXiv manuscript is refreshed quarterly as newly accepted tasks are incorporated.
Still have questions?
Email us at sunyiyou@berkeley.edu or hanxinyang@berkeley.edu.