Megaton

Megaton Video Bench Evaluation v2.1

How Video Index scores are calculated

Megaton combines consistent test conditions, prompt-level evidence, calibrated professional judgment, and versioned calculations to compare generative video models transparently.

Evaluation process

01

Standardized inputs

Models are evaluated against the same versioned prompt corpus. We record the provider, model revision, generation settings, output, attempts, and collection date so every result retains its operating context.

02

Observable evidence

Reviewers score what appears in the delivered video. Prompt-specific requirements, category rubrics, evidence frames, technical signals, and failure notes remain attached to the decision instead of being collapsed into an unexplained opinion.

03

Comparable scoring

Deterministic checklists award 10 points for a fulfilled requirement, 5 for a partial result, and 0 when it is missing. Expert-judged categories use calibrated 1–10 anchors and convert them to the same 0–100 scale.

04

Versioned publication

Category results stay separate and auditable. Published indices use one locked benchmark edition and formula, preserve missing or excluded coverage explicitly, and never blend scores from incompatible benchmark versions.

The benchmark edition, scoring rubric, and source evidence are preserved together so a score can be reviewed in context rather than treated as a permanent property of a model.

View the leaderboard