AGI Research
Agentic AI: evaluation, alignment, and autonomy.
The research bench for agentic AI. Two pillars: how we measure agent behavior, and how the agents themselves are built and improved.
Evaluation and Alignment The trace evaluation standard, the failure taxonomy, calibration against the human vote, and the judge.
Autonomous Agents What the agent stack has to emit to be measurable at all, and how its task guides are generated and served.