Skip to content

Benchmark methodology

Official comparisons freeze challenge versions, grader versions, environment images, and model configuration. Historical rows are never rewritten when a challenge changes.

One run does not rank a model. Official tracks require a minimum sample size. Medians, variance, and infrastructure-failure exclusions will be published with each named release (for example DE-2026.1).

Temperature and tool limits are recorded per run. Provider outages are classified separately from solution failure and do not penalize agent score.

Official catalog content is DE-Core-2026.1 (30 challenges). Workspace Solve myself, Run with agent, and Compare models paths execute isolated fixtures and write hidden-test scores. The public Agent Index is still a platform sample, not a published official rank.