A smaller context is not automatically a better result. A reduction matters only if the agent can still make the right change, preserve behavior, and finish the task with acceptable latency and overhead.
A useful evaluation would compare matched repository tasks with and without Atrium under the same model, task instructions, and execution budget. It should report task success, regression rate, tokens, latency, indexing cost, and failure categories, alongside repository and environment details.
01Context costTokens and repeated input
02Task qualityCorrectness and completion
03System overheadLatency, indexing, maintenance
No performance advantage is claimed on this page. Publish a number only when the task set, baseline, model configuration, measurement process, and limitations can be inspected and repeated.
Read the evaluation method ↗