<report>

Alibaba

Wan 3.0

v-benchmark v2 by Megaton benchmark
Evaluated · September 1, 2026Leaderboard rank #2 of 14Last updated · September 1, 2026

81.62

/ 100

Categories

C01

Prompt AdherenceWhether every explicit, observable instruction in the prompt is fulfilled.

This Video

95.5

This Video · Assessment

This video fulfills its explicit, observable instructions with near-complete consistency.

This Category

94.7

42 videos were rated for this category

This Category · Assessment

Across the evaluated prompts, explicit subjects, actions, relationships, negations, and required text are fulfilled with near-complete consistency.

Models that scored similarly on this category

Violations

This Video
Essential1
Major2
Minor3
This Category
Essential9
Major10
Minor11

42 videos

C02

Scene ConsistencyWhether established subjects, objects, spatial relationships, environment, and lighting remain coherent over time.

This Video

69.7

This Video · Assessment

This video maintains subjects, objects, spatial relationships, environment, and lighting over time in a generally coherent way, though visible failures recur.

This Category

92.9

36 videos were rated for this category

This Category · Assessment

Subjects, objects, spatial relationships, environment, and lighting remain coherent across nearly all evaluated videos.

Models that scored similarly on this category

Violations

This Video
Essential5
Major3
Minor0
This Category
Essential10
Major7
Minor2

36 videos

C03

PhysicsWhether the central motion, force, material, and reaction behave coherently in the requested world and style.

This Video

80.0

This Video · Assessment

Convincing overall, with a few visible but non-disruptive physical anomalies.

This Category

71.8

17 videos were rated for this category

This Category · Assessment

Motion and interaction are convincing overall, with recurring but non-disruptive physical anomalies.

Models that scored similarly on this category

C04

Human FidelityWhether faces, eyes, skin, bodies, hands, motion, and human presence remain anatomically and perceptually convincing.

This Video

100.0

This Video · Assessment

Fully convincing within the requested style, with no meaningful uncanny or anatomical defects.

This Category

86.7

27 videos were rated for this category

This Category · Assessment

Human rendering is highly convincing across the set, with only small localized defects.

Models that scored similarly on this category

C05

2D Animation Craft & ArtistryThe quality and artistic authorship of intentionally 2D animation within its requested conventions.

This Video

80.0

This Video · Assessment

Technically excellent and fully professional, but more conventional or less expressive than the highest tier.

This Category

70.0

12 videos were rated for this category

This Category · Assessment

The work is strongly and convincingly 2D overall, though execution or authorship is less consistent than the highest tier.

Models that scored similarly on this category

C06

3D Animation Craft & ArtistryThe quality, coherence, and artistic authorship of intentionally 3D animation.

This Video

70.0

This Video · Assessment

Strong and convincing, with visible craft or artistic limitations.

This Category

76.0

5 videos were rated for this category

This Category · Assessment

The work is technically strong and convincing overall, though craft or artistic limitations recur.

Models that scored similarly on this category

C07

Cinematography Craft & ArtistryShot design, framing, camera movement, lighting, focus, visual hierarchy, and intentional filmmaking authorship.

This Video

80.0

This Video · Assessment

Technically excellent and fully professional, but more conventional or less expressive.

This Category

72.9

7 videos were rated for this category

This Category · Assessment

Shot design is technically strong overall, though recurring choices are more conventional or less expressive.

Models that scored similarly on this category

C08

Taste & Art DirectionThe strength, specificity, coherence, and artistic authorship of the result as a finished creative choice.

This Video

60.0

This Video · Assessment

Competent but visibly reliant on familiar AI aesthetics or underdeveloped art direction.

This Category

68.1

60 videos were rated for this category

This Category · Assessment

Art direction is generally capable, but generic, unresolved, or less-confident choices recur.

Models that scored similarly on this category

C09

Animal & Creature FidelityWhether real animals and designed creatures preserve convincing anatomy, identity, motion, behavior, and interaction.

This Video

70.0

This Video · Assessment

Believable overall, with clear defects under inspection.

This Category

84.3

7 videos were rated for this category

This Category · Assessment

Animal and creature fidelity is highly convincing across the set, with only small localized defects.

Models that scored similarly on this category

C10

Object & Product FidelityWhether requested objects have correct identity, structure, materials, details, motion, and interactions.

This Video

100.0

This Video · Assessment

This video preserves the requested objects’ identity, structure, materials, details, motion, and interactions with near-complete consistency.

This Category

98.7

8 videos were rated for this category

This Category · Assessment

Requested objects preserve identity, structure, materials, details, motion, and interactions with near-complete consistency.

Models that scored similarly on this category

Violations

This Video
Essential0
Major3
Minor0
This Category
Essential0
Major3
Minor0

8 videos

C11

Causal & Semantic CoherenceWhether events, relationships, transformations, and cause-and-effect remain semantically coherent, including surreal premises.

This Video

70.0

This Video · Assessment

The event remains understandable, but noticeable contradictions appear.

This Category

78.2

11 videos were rated for this category

This Category · Assessment

Events remain clear overall, though recurring logical irregularities are visible but bounded.

Models that scored similarly on this category

C12

Text FidelityWhether requested text content, typography, placement, attachment, legibility, and temporal stability are correct.

This Video

86.1

This Video · Assessment

This video preserves requested text content, typography, placement, attachment, legibility, and temporal stability very reliably, with only isolated limitations.

This Category

82.5

6 videos were rated for this category

This Category · Assessment

Text fidelity is highly reliable, with only isolated content, placement, or temporal defects.

Models that scored similarly on this category

Violations

This Video
Essential1
Major0
Minor0
This Category
Essential4
Major3
Minor0

6 videos

C13

Character Performance & ActingWhether character intention, emotion, timing, reactions, eyelines, gesture, and performance read convincingly.

This Video

70.0

This Video · Assessment

Believable overall, with visible limitations in timing, expression, or specificity.

This Category

66.0

21 videos were rated for this category

This Category · Assessment

Acting is generally believable, but synthetic timing, expression, attention, or specificity recurs.

Models that scored similarly on this category

Differences between other models

CategoryWan 3.0Seedance 2.5Seedance 2.0FLUX 3Gemini Omni 1.1 Flash
C01Prompt Adherence94.787.887.487.187.2
C03Physics71.884.166.570.658.2
C02Scene Consistency92.997.292.294.888.8
C04Human Fidelity86.793.382.273.787.0
C052D Animation Craft & Artistry70.080.066.760.464.6
C063D Animation Craft & Artistry76.076.070.066.076.0
C11Causal & Semantic Coherence78.293.680.980.972.7
C08Taste & Art Direction68.175.063.265.065.8
C07Cinematography Craft & Artistry72.980.068.671.468.6
C09Animal & Creature Fidelity84.395.780.075.787.1
C10Object & Product Fidelity98.7100.098.283.899.1
C12Text Fidelity82.593.385.882.492.1
C13Character Performance & Acting66.075.066.768.074.0

Leaderboard

Methodology

The Megaton Index is the published 0–100 weighted summary for v-benchmark v2. This public report shows the model's Index and category scores; the full calibration evidence, decision records, and scoring implementation are reserved for the full report.

Read the full methodology →

Changelog

Published
September 1, 2026
V-Benchmark version
2.0.0
Megaton Index
81.62
Published rank
2 of 14
View model changelog →