METHODS & EVIDENCE

Methods that make claims
checkable.

A lightweight but explicit approach to implementation, evaluation, and reporting. The aim is not to make every experiment look conclusive; it is to make each claim easier to inspect and challenge.

IDEACTIVATE / WORKING METHODHYPOTHESES ARE NOT RESULTS

Research claims become useful when another person can understand the setup, see the trade-offs, and attempt to repeat the result. Ideactivate is small, so the process needs to be lightweight—but it still needs to be explicit.

1. Define the question before the metric.

Start by stating the behavior the system should improve. For Atrium, that means a repository-level task outcome, not token reduction in isolation. For EILER, it means a measurable sequence-model behavior, not architectural novelty. For FluxCut, it means a reproducible edit or media operation, not a screenshot.

2. Separate implementation from hypothesis.

Project pages distinguish current implementation, intended architecture, ongoing work, and unvalidated ideas. A diagram explains a proposed system; it does not establish that the implementation is complete or that it works as intended.

3. Compare against an appropriate baseline.

A useful evaluation holds relevant conditions constant and makes the comparison inspectable. When a model or coding agent is involved, document the model version, task set, instructions, tool configuration, environment, and budget. Use ablations when they help isolate the contribution of one component.

4. Report the costs with the result.

For developer tools, include task completion, regressions, tokens, latency, indexing or setup overhead, and failure categories. For sequence models, include data and tokenization, training budget, parameter count, memory use, throughput, convergence, and variability where relevant. For desktop software, test state transitions, timing, recovery, and error behavior.

5. Treat failures as data.

Unsupported syntax, stale context, invalid memory updates, numerical instability, and media edge cases are part of the problem definition. A result that only describes the happy path gives readers too little information to decide where an approach is useful.

6. Make the evidence easy to inspect.

As artifacts mature, this site will link to the corresponding source code, commands, test fixtures, and result files. Until there is a reproducible report, the page will describe the evaluation plan rather than publish an unverified percentage.

Current evaluation priorities

Atrium: matched repository tasks with and without structured context; task success, regression rate, tool overhead, tokens, latency, and index freshness.

EILER: well-tuned reference models, ablations of the state and memory paths, controlled training budgets, long-range information tasks, and resource costs.

FluxCut: deterministic timeline operations, repeatable undo/redo, frame and audio timing, missing-asset recovery, and export consistency.

RELATED PAGESResearch agenda ↗Atrium evaluation plan ↗Current priorities ↗