01
Standardized inputs
Models are evaluated against the same versioned prompt corpus. We record the provider, model revision, generation settings, output, attempts, and collection date so every result retains its operating context.
Megaton Video Bench Evaluation v2.1
Megaton combines consistent test conditions, prompt-level evidence, calibrated professional judgment, and versioned calculations to compare generative video models transparently.
01
Models are evaluated against the same versioned prompt corpus. We record the provider, model revision, generation settings, output, attempts, and collection date so every result retains its operating context.
02
Reviewers score what appears in the delivered video. Prompt-specific requirements, category rubrics, evidence frames, technical signals, and failure notes remain attached to the decision instead of being collapsed into an unexplained opinion.
03
Deterministic checklists award 10 points for a fulfilled requirement, 5 for a partial result, and 0 when it is missing. Expert-judged categories use calibrated 1–10 anchors and convert them to the same 0–100 scale.
04
Category results stay separate and auditable. Published indices use one locked benchmark edition and formula, preserve missing or excluded coverage explicitly, and never blend scores from incompatible benchmark versions.
The benchmark edition, scoring rubric, and source evidence are preserved together so a score can be reviewed in context rather than treated as a permanent property of a model.
View the leaderboard