AI Calculation Benchmark

A planned, reproducible evaluation of formula derivation, unit safety, boundary handling, and source support across AI models.

Protocol preview. The study has not produced a complete public result set or model leaderboard. No performance claim should be inferred from this page.

Research question

When multiple current AI models receive the same frozen source material and calculation task, how often do they independently derive the same valid formula, preserve units and domain boundaries, and support factual claims with the supplied sources?

Planned protocol

  1. Freeze the task set

    Select standard and maintained-risk capabilities from the public registry. Publish IDs, source snapshots, expected units, and deterministic oracle tests before scoring models.

  2. Separate model roles

    Model A drafts a derivation, Model B independently recomputes without seeing A, and Model C performs an adversarial review for unit, boundary, citation, and unsupported-claim errors. Exact provider, model version, date, parameters, and prompts will be reported.

  3. Score with code

    Programmatic checks compare outputs with independent fixtures and property tests. Model agreement alone cannot override a failing deterministic oracle or a prohibited risk class.

  4. Publish reproducibility artifacts

    A complete release must include the frozen corpus manifest, prompt and model versions, run identifiers, scoring code, per-task outcomes, aggregate results, and known limitations.

Metrics defined for the first release

Formula agreement
Equivalent derivation after normalization, confirmed against an independent oracle.
Unit safety
Correct dimensional treatment, conversions, output units, and incompatible-unit rejection.
Boundary handling
Correct behavior for zero, negative, empty, singular, and declared domain-limit cases.
Source support
Material factual claims trace to the frozen source set without invented attribution.
Unsupported assertion rate
Share of evaluated claims that exceed or contradict the supplied evidence.

Safety and interpretation boundaries

Browse the Open Formula Registry for the current public capability metadata.