AI Calculation Benchmark
A planned, reproducible evaluation of formula derivation, unit safety, boundary handling, and source support across AI models.
Research question
When multiple current AI models receive the same frozen source material and calculation task, how often do they independently derive the same valid formula, preserve units and domain boundaries, and support factual claims with the supplied sources?
Planned protocol
Freeze the task set
Select standard and maintained-risk capabilities from the public registry. Publish IDs, source snapshots, expected units, and deterministic oracle tests before scoring models.
Separate model roles
Model A drafts a derivation, Model B independently recomputes without seeing A, and Model C performs an adversarial review for unit, boundary, citation, and unsupported-claim errors. Exact provider, model version, date, parameters, and prompts will be reported.
Score with code
Programmatic checks compare outputs with independent fixtures and property tests. Model agreement alone cannot override a failing deterministic oracle or a prohibited risk class.
Publish reproducibility artifacts
A complete release must include the frozen corpus manifest, prompt and model versions, run identifiers, scoring code, per-task outcomes, aggregate results, and known limitations.
Metrics defined for the first release
- Formula agreement
- Equivalent derivation after normalization, confirmed against an independent oracle.
- Unit safety
- Correct dimensional treatment, conversions, output units, and incompatible-unit rejection.
- Boundary handling
- Correct behavior for zero, negative, empty, singular, and declared domain-limit cases.
- Source support
- Material factual claims trace to the frozen source set without invented attribution.
- Unsupported assertion rate
- Share of evaluated claims that exceed or contradict the supplied evidence.
Safety and interpretation boundaries
- No human expert currently reviews or certifies this benchmark.
- High-risk medical, emergency, poisoning, and professional-safety decisions are excluded.
- A high benchmark score would measure the declared task set, not general authority.
- Model names and scores will remain absent until the required artifacts can be published.
Browse the Open Formula Registry for the current public capability metadata.