Checklists
Evaluators score a defined set of requirements. Missing essential or major requirements can cap the result at 40% or 60%.
Swipe for more top models

V-Benchmark measures how well video models follow a prompt, represent motion and physics, maintain a coherent scene, and produce usable visual results. We combine expert review, automated auditing, and repeated technical checks across 15 categories.
Last updated ·
Methodology version · 2.1.0
View methodology changelog →We test each model with a fixed set of prompts. Film and broadcast professionals score one output per prompt against checklists and category-specific rubrics. Evaluators do not see the model or provider name, and reference examples help them apply the same standard.
Technical measurements check eligible clips for problems such as visual instability, frozen frames, exposure, compression, and frame-rate issues. Multimodal models audit the results and flag possible errors for human review; they do not change scores automatically.
Each report publishes category scores so readers can see a model’s strengths and weaknesses. Weighted category scores produce the Megaton Index for broad comparison. V-Benchmark complements preference-based arenas, which ask viewers to choose between outputs rather than score one model against defined criteria.
Trained evaluators from film and broadcast score most categories. We use four methods to match the type of evidence each category requires:
Evaluators score a defined set of requirements. Missing essential or major requirements can cap the result at 40% or 60%.
Evaluators map category-specific criteria to a 1–10 score and compare their judgment with approved reference examples.
Combines a checklist with a scoring rubric.
Uses repeatable technical measurements that a human reviews before publication.
| Category | Definition | Mode | Weight |
|---|---|---|---|
| C01Prompt Adherence | Whether every instruction in the prompt is fulfilled. | Checklist | 20% |
| C02Scene Consistency | Whether established subjects, objects, spatial relationships, environment, and lighting remain consistent over time. | Mixed | 10% |
| C03Physics | Whether physical interactions & forces mirror real life. | Mixed | 20% |
| C04Human Fidelity | Whether human features and behaviors are rendered accurately and consistently. | Mixed | 10% |
| C052D Animation Craft & Artistry | The quality and artistic skill of 2D animation given the conventions of the craft | Scoring Rubric | 7.5% |
| C063D Animation Craft & Artistry | The quality and artistic skill of 3D animation given the conventions of the craft | Scoring Rubric | 7.5% |
| C07Cinematography Craft & Artistry | The model’s aptitude at shot design, composition, camera movement, lighting, visual hierarchy, and filmmaking fluency. | Scoring Rubric | 2.5% |
| C08Taste & Art Direction | The artistic strength & competence exhibited by the model | Scoring Rubric | 6.5% |
| C09Animal & Creature Fidelity | Measures whether animal and creatures demonstrate convincing anatomy, preserve their identity, move realistically, behave naturally and interact as expected. | Mixed | 1.5% |
| C10Object & Product Fidelity | Measures whether an object is identifiable, maintains it’s structure, is rendered in its expected material and maintains its form given the context of the scene. | Mixed | 1.5% |
| C11Causal & Semantic Coherence | Whether events, relationships, transformations, and cause-and-effect remain semantically coherent. | Mixed | 7.5% |
| C12Text Fidelity | Whether requested text is rendered legibly and remains coherent throughout a generation | Mixed | 1.5% |
| C13Character Performance & Acting | The skill and ability of human characters to perform convincingly. | Scoring Rubric | 1.5% |
| C14Delivered File & Playback Integrity | Measures if a delivered file plays correctly without freezes or encoding issues. Includes whether files are in the requested format, length and frame rate. | Reviewed measurement | 0.5% |
| C15Quality + Reliability | A video’s stability and reliability as determined by a suite of automated checks. | Reviewed measurement | 2% |
| Publication weights total 100% | 100% | ||
V-Benchmark weights are determined through expert judgment and are intended to represent a helpful but limited general performance score best used to compare models to one another.
Greater weight is assigned to categories where failures can make an output fundamentally unusable or materially fail a user's request, including prompt adherence, physics, scene consistency as well as coherence. As such these categories tend to have broader prompt coverage and criteria that limits judgement discretion.
Stylistic, artistic or craft categories receive lower weights given their potential for disagreement among expert judges. Although these capabilities are important to judge in a creative model they are more dependent on aesthetic or craft judgment or they may apply to a narrower portion of the generated set. They’re lower weighting reflects an attempt to mitigate possible variance in judgement
These weights reflect a deliberate balance between how severely a model fails our confidence that the measure is reproducible as well as if we determine that a score could elicit reasonable pushback from an experienced and trained peer.
Category scores are therefore included and we encourage readers to review them alongside the index in order to discern whether a model maybe better suited for their own use cases
The 2D Animation Index, Physics Index and Prompt Adherence Index refer to the corresponding V-Benchmark category scores. They are category-level views rather than separate composite scores.
V-benchmark was developed with an initial evaluation suite of 60 prompts (Full) for our highest weighted categories (Physics 20% weight & Prompt adherence 20% weight and scene consistency 10% weight) as well as certain qualitative categories like Taste. Given our aim to evaluate all models and evaluate them within a reasonable timeframe, video models may be scored on a 36 (Core) or 48 prompt suite (Core+). Scores that are run on more limited coverage runs are marked and readers are encouraged to be mindful of the error confidence interval of +/- 2 that they produce.
36-prompt run
Error confidence interval (+/- 2 Index points)
A concise evaluation designed to estimate the final score on the highest weight categories and specific qualitative categories.
48-prompt cap
Error confidence interval (+/- 2 Index points)
A larger intermediate panel that reduces error confidence in the final index score.
60-prompt cap
Error confidence interval (Full reference)
The complete run for every public category used for the standard full publication.
Up to 3,000 generations
Error confidence interval (Study-specific)
A commissioned study that can add repeats, variants, common failure modes, supported settings, new and different modalities, and targeted diagnostics while retaining a comparable public scorecard.
These are private evaluations run on behalf of labs & companies. Please email founders@megaton.ai to inquire about our extensive evaluation services.
To get a sense of how accurate our 36 and 48 run evaluations are compared to our full 60 prompt evaluations we tested three full evaluations to simulate how close a 36 and 48 run evaluation would have fared to arrive at our final index number. We ran these using the Seedance 2.5, Seedance 2.0, and MiniMax Hailuo 03 runs whose full report you can find under our evaluations page.
Across this group we observed that the 36 prompt Core run only differs from the Full score by an average of 0.23 points with a maximum observed difference of 0.40 points/ The 48 prompt run (Core+) averages a 0.08 difference with a max change of 0.18.
We also ran 300,000 stratified resamples per model at each eval depth with the following results:
The 95th percentile deviation was 0.80 points for Core and 0.50 points for Core+.
Maximum observed deviations were 1.65 and 0.95 points respectively. All 300,000 resamples at each depth remained within ±2 Index points of the Full score.
We therefore use ±2 Index points as a conservative empirical comparison for Core & Core+ evaluations.
To be clear, this only describes the variation observed with a limited and should not be interpreted as a universal statistical confidence interval.
We will continue to run this experiment as more models are evaluated using our full prompt suite and will update our methodology and changelog accordingly.
| Depth | Mean absolute difference | Fixed-panel maximum | Resampled 95th percentile | Resampled maximum |
|---|---|---|---|---|
| Core | 0.23 | 0.40 | 0.80 | 1.65 |
| Core+ | 0.08 | 0.18 | 0.50 | 0.95 |
API provider date and time are reported on model scores.
If prompts are rejected by provider guardrails we may substitute a similar prompt. When done so we disclose in our reporting.
Core, Core+ and Full evaluations use trained evaluators to score each item. Evaluators are drawn from the film and broadcast industry and trained on V-Benchmark criteria, reference examples, as well as evaluation procedures & methedology.
Evaluators are blind to model and provider while scoring while the same evaluation criteria and standards are applied across models to reduce variation in scoring between any two judges.
Each prompt produces one generation in Core, Core+, and Full evaluations. V-Benchmark therefore measures performance across a defined set of individual generation attempts rather than selecting the best result from multiple samples.
Each generation is scored between 1 to 10. This is then multiplied by 10 and used to produce a 10-100 score to be used to weight our category score. This 100 point score is what’s used in our final index calculation.
For checklists each item is classified by its level of importance. These are “Essential”, “Major”, “Minor”. This classification weighs its impact on the 1 to 10 score with 10 = all requirements fulfilled and 1 = no requirements fulfilled. The weighting however means that if an essential or major element is missing a score will be weighted down proportionally according to the specific weight it’s assigned. Some checklist items also cap the score to prevent a major / essential issue from being averaged away.
Categories judged through Scoring rubrics are done through a predetermined criteria. We use historical scores as reference points for each score from 1-10 to ensure judges are maintaining a consistent standard while judging
Multimodal models are used to audit the final result and flag potential inconsistencies, missed failures or potential scoring errors. These audits are flagged for additional human inspection prior to publication but do not automatically override the evaluator’s score.
This automated audit is calibrated against historical human-scored evaluation data to identify results that materially disagree with expected scoring patterns.
We supplement expert evaluation with a technical analysis of every eligible delivered clip. These measurements evaluate reproducible properties of the video file for reporting and quickly flagging technical deficiencies in the delivered file.
We measure this under our Deterministic Fidelity category which produces a 0–100 score split equally between quality and reliability.
Quality measures motion-compensated visual stability using three equally weighted measures:
Reliability measures the percentage of applicable clips passing nine technical checks:
Quality and Reliability each contribute 50% of the Deterministic Fidelity score. Deterministic Fidelity contributes 2.5% of the final Megaton Index.
Motion Integrity reports the average of the three Quality measures, while Technical Cleanliness reports the percentage of clips passing the reliability test. These are diagnostic summaries and do not receive additional weight.
We also report characteristics including color-temperature stability, color usage, fine-detail level, motion degree, native frame rate, and tonal range which we report descriptively and do not affect the Megaton Index.
Megaton earns revenue from commission evaluation services, advisory & consulting services, advertising, sponsorships, paid links, and affiliate links. Clients and commercial partners do not determine the evaluation scores, rankings, findings, or editorial conclusions Megaton publishes.