mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
feat(phase-12/17): video-language models and temporal grounding
This commit is contained in:
+84
@@ -0,0 +1,84 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 520" font-family="Georgia, 'Times New Roman', serif">
|
||||
<defs>
|
||||
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
|
||||
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
|
||||
</marker>
|
||||
<style>
|
||||
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
|
||||
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
|
||||
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
|
||||
.reg { fill: #e9e6ff; stroke: #5a4fcf; stroke-width: 1.5; }
|
||||
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
|
||||
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
|
||||
.caption { font-size: 11px; fill: #555; font-style: italic; }
|
||||
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
|
||||
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
|
||||
</style>
|
||||
</defs>
|
||||
|
||||
<text x="480" y="24" text-anchor="middle" class="title">Video VLMs — frame sampling, temporal tokens, grounding output</text>
|
||||
|
||||
<rect x="30" y="50" width="900" height="220" class="box"/>
|
||||
<text x="480" y="72" text-anchor="middle" class="head">three architecture patterns from 2023 to 2025</text>
|
||||
|
||||
<rect x="50" y="90" width="280" height="160" class="hot"/>
|
||||
<text x="190" y="110" text-anchor="middle" class="step">Video-LLaMA (2023)</text>
|
||||
<text x="190" y="128" text-anchor="middle" class="small">Q-former + audio branch</text>
|
||||
<text x="190" y="144" text-anchor="middle" class="small">16 frames @ 2 FPS fixed</text>
|
||||
<text x="190" y="160" text-anchor="middle" class="small">32 video queries, 32 audio</text>
|
||||
<text x="190" y="184" text-anchor="middle" class="step">strength: audio grounding</text>
|
||||
<text x="190" y="204" text-anchor="middle" class="small">weakness: 8s fixed clip</text>
|
||||
<text x="190" y="220" text-anchor="middle" class="small">no event time localization</text>
|
||||
|
||||
<rect x="340" y="90" width="280" height="160" class="cool"/>
|
||||
<text x="480" y="110" text-anchor="middle" class="step">Video-LLaVA (2023)</text>
|
||||
<text x="480" y="128" text-anchor="middle" class="small">MLP + shared encoder</text>
|
||||
<text x="480" y="144" text-anchor="middle" class="small">8 frames @ 1-2 FPS</text>
|
||||
<text x="480" y="160" text-anchor="middle" class="small">alignment before projection</text>
|
||||
<text x="480" y="184" text-anchor="middle" class="step">strength: simple + effective</text>
|
||||
<text x="480" y="204" text-anchor="middle" class="small">weakness: short clips only</text>
|
||||
<text x="480" y="220" text-anchor="middle" class="small">no dynamic FPS</text>
|
||||
|
||||
<rect x="630" y="90" width="280" height="160" class="cold"/>
|
||||
<text x="770" y="110" text-anchor="middle" class="step">Qwen2.5-VL (2025)</text>
|
||||
<text x="770" y="128" text-anchor="middle" class="small">TMRoPE + dynamic FPS</text>
|
||||
<text x="770" y="144" text-anchor="middle" class="small">arbitrary duration</text>
|
||||
<text x="770" y="160" text-anchor="middle" class="small">absolute time tokens</text>
|
||||
<text x="770" y="184" text-anchor="middle" class="step">strength: event grounding</text>
|
||||
<text x="770" y="204" text-anchor="middle" class="small">JSON output format</text>
|
||||
<text x="770" y="220" text-anchor="middle" class="small">open SOTA 2026</text>
|
||||
|
||||
<rect x="30" y="290" width="900" height="230" class="box"/>
|
||||
<text x="480" y="312" text-anchor="middle" class="head">frame sampling + output format</text>
|
||||
|
||||
<rect x="60" y="330" width="260" height="170" class="reg"/>
|
||||
<text x="190" y="352" text-anchor="middle" class="step">frame sampling strategies</text>
|
||||
<text x="190" y="372" text-anchor="middle" class="small">uniform: N frames / duration</text>
|
||||
<text x="190" y="388" text-anchor="middle" class="small">- simple, loses motion peaks</text>
|
||||
<text x="190" y="408" text-anchor="middle" class="small">dynamic FPS: motion-weighted</text>
|
||||
<text x="190" y="424" text-anchor="middle" class="small">- denser in high-motion spans</text>
|
||||
<text x="190" y="444" text-anchor="middle" class="small">event-driven: detector + sample</text>
|
||||
<text x="190" y="460" text-anchor="middle" class="small">- best for action recognition</text>
|
||||
<text x="190" y="480" text-anchor="middle" class="caption">pair with 3x3 pooling per frame</text>
|
||||
|
||||
<rect x="340" y="330" width="280" height="170" class="hot"/>
|
||||
<text x="480" y="352" text-anchor="middle" class="step">grounding output formats</text>
|
||||
<text x="480" y="372" text-anchor="middle" class="small">free text:</text>
|
||||
<text x="480" y="388" text-anchor="middle" class="small">"The cat jumps around 4s"</text>
|
||||
<text x="480" y="412" text-anchor="middle" class="small">JSON:</text>
|
||||
<text x="480" y="428" text-anchor="middle" class="small">{"event":"jump","start":4.1,"end":4.3}</text>
|
||||
<text x="480" y="452" text-anchor="middle" class="small">token:</text>
|
||||
<text x="480" y="468" text-anchor="middle" class="small">"<time>4.1</time> jump"</text>
|
||||
<text x="480" y="488" text-anchor="middle" class="caption">JSON is easiest to parse downstream</text>
|
||||
|
||||
<rect x="640" y="330" width="280" height="170" class="cool"/>
|
||||
<text x="780" y="352" text-anchor="middle" class="step">benchmarks</text>
|
||||
<text x="780" y="372" text-anchor="middle" class="small">VideoMME: general, 2500 samples</text>
|
||||
<text x="780" y="388" text-anchor="middle" class="small">TempCompass: before/after</text>
|
||||
<text x="780" y="404" text-anchor="middle" class="small">EgoSchema: 3min first-person</text>
|
||||
<text x="780" y="420" text-anchor="middle" class="small">Video-MMMU: multi-discipline</text>
|
||||
<text x="780" y="446" text-anchor="middle" class="step">open SOTA 2026</text>
|
||||
<text x="780" y="462" text-anchor="middle" class="small">Qwen2.5-VL-72B</text>
|
||||
<text x="780" y="478" text-anchor="middle" class="small">TMRoPE is the differentiator</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 5.8 KiB |
@@ -0,0 +1,141 @@
|
||||
"""Video VLM frame sampler + temporal-grounding evaluator — stdlib.
|
||||
|
||||
Three toys:
|
||||
1. Uniform frame sampler.
|
||||
2. Dynamic-FPS sampler using motion proxy (synthetic per-frame motion scalar).
|
||||
3. Temporal-grounding evaluator with IoU-style scoring.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
import random
|
||||
from dataclasses import dataclass
|
||||
|
||||
random.seed(4)
|
||||
|
||||
|
||||
def uniform_sample(duration: float, n: int) -> list[float]:
|
||||
if n <= 1:
|
||||
return [duration / 2]
|
||||
step = duration / n
|
||||
return [round(step * (i + 0.5), 3) for i in range(n)]
|
||||
|
||||
|
||||
def dynamic_sample(motion: list[float], fps_cap: int = 4,
|
||||
total_budget: int = 32) -> list[float]:
|
||||
"""Allocate samples by per-second motion; cap per second at fps_cap."""
|
||||
total_motion = sum(motion)
|
||||
if total_motion == 0:
|
||||
return uniform_sample(len(motion), total_budget)
|
||||
samples_per_sec = []
|
||||
for m in motion:
|
||||
raw = total_budget * m / total_motion
|
||||
samples_per_sec.append(min(fps_cap, max(1, round(raw))))
|
||||
times = []
|
||||
for sec_idx, count in enumerate(samples_per_sec):
|
||||
for j in range(count):
|
||||
t = sec_idx + (j + 0.5) / count
|
||||
times.append(round(t, 3))
|
||||
return times
|
||||
|
||||
|
||||
def iou(a_start: float, a_end: float, b_start: float, b_end: float) -> float:
|
||||
inter = max(0.0, min(a_end, b_end) - max(a_start, b_start))
|
||||
union = max(a_end, b_end) - min(a_start, b_start)
|
||||
return inter / union if union > 0 else 0.0
|
||||
|
||||
|
||||
@dataclass
|
||||
class Event:
|
||||
name: str
|
||||
start: float
|
||||
end: float
|
||||
|
||||
|
||||
def evaluate_grounding(predictions: list[Event], ground_truth: list[Event],
|
||||
tol_iou: float = 0.3) -> dict:
|
||||
hits = 0
|
||||
details = []
|
||||
for gt in ground_truth:
|
||||
best_iou = 0.0
|
||||
best_pred = None
|
||||
for p in predictions:
|
||||
if p.name == gt.name:
|
||||
val = iou(p.start, p.end, gt.start, gt.end)
|
||||
if val > best_iou:
|
||||
best_iou = val
|
||||
best_pred = p
|
||||
hit = best_iou >= tol_iou
|
||||
if hit:
|
||||
hits += 1
|
||||
details.append((gt.name, best_iou, hit))
|
||||
return {"recall": hits / max(1, len(ground_truth)), "details": details}
|
||||
|
||||
|
||||
def demo_samplers() -> None:
|
||||
print("\nFRAME SAMPLING STRATEGIES")
|
||||
print("-" * 60)
|
||||
duration = 10.0
|
||||
uni = uniform_sample(duration, 8)
|
||||
print(f" uniform (8 frames / 10s) : {uni}")
|
||||
motion = [0.1, 0.1, 0.8, 0.9, 0.9, 0.2, 0.1, 0.5, 0.9, 0.9]
|
||||
dyn = dynamic_sample(motion, fps_cap=4, total_budget=12)
|
||||
print(f" motion : {motion}")
|
||||
print(f" dynamic (12 frames total): {dyn}")
|
||||
print(" dynamic places more frames in high-motion seconds 2-4 and 7-9")
|
||||
|
||||
|
||||
def demo_grounding() -> None:
|
||||
print("\nTEMPORAL GROUNDING EVAL (IoU >= 0.3)")
|
||||
print("-" * 60)
|
||||
ground = [
|
||||
Event("jump", 4.0, 4.5),
|
||||
Event("turn", 6.0, 6.5),
|
||||
Event("sit", 8.5, 9.5),
|
||||
]
|
||||
predictions = [
|
||||
Event("jump", 4.1, 4.7),
|
||||
Event("turn", 5.8, 6.2),
|
||||
Event("sit", 9.2, 9.6),
|
||||
]
|
||||
result = evaluate_grounding(predictions, ground)
|
||||
print(f" recall@IoU0.3 : {result['recall']:.2f}")
|
||||
for name, val, hit in result["details"]:
|
||||
tag = "HIT" if hit else "miss"
|
||||
print(f" {name:<6} IoU={val:.2f} {tag}")
|
||||
|
||||
|
||||
def arch_compare() -> None:
|
||||
print("\nVIDEO VLM ARCHITECTURES")
|
||||
print("-" * 60)
|
||||
rows = [
|
||||
("Video-LLaMA", "Q-former / 16 frames", "fixed clip, audio branch"),
|
||||
("Video-LLaVA", "MLP / 8 frames", "shared image+video encoder"),
|
||||
("VILA-1.5", "MLP / 8-16 frames", "pretraining-heavy"),
|
||||
("Qwen2.5-VL", "TMRoPE / dynamic FPS", "absolute time, best open 2025"),
|
||||
("LLaVA-OV-1.5", "pool / 32 frames", "unified image+multi+video"),
|
||||
]
|
||||
print(f" {'model':<14}{'compressor':<24}{'note'}")
|
||||
for r in rows:
|
||||
print(f" {r[0]:<14}{r[1]:<24}{r[2]}")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("=" * 60)
|
||||
print("VIDEO-LANGUAGE TEMPORAL GROUNDING (Phase 12, Lesson 17)")
|
||||
print("=" * 60)
|
||||
|
||||
demo_samplers()
|
||||
demo_grounding()
|
||||
arch_compare()
|
||||
|
||||
print("\nTAKEAWAY")
|
||||
print("-" * 60)
|
||||
print(" temporal tokens matter as much as the visual encoder")
|
||||
print(" dynamic FPS + TMRoPE is the 2026 open-source default")
|
||||
print(" JSON grounded output beats free-text for downstream use")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,149 @@
|
||||
# Video-Language Models: Temporal Tokens and Grounding
|
||||
|
||||
> Video is not a stack of photos. A 5-second clip has causal ordering, action verbs, and event timing that an image model cannot represent. Video-LLaMA (Zhang et al., June 2023) shipped the first open video-LLM with audio-visual grounding. VideoChat and Video-LLaVA scaled the pattern. By 2025 Qwen2.5-VL's TMRoPE closed the gap with frontier proprietary models. Each system solved temporal tokens differently — Q-former per clip, concat-pool per frame, TMRoPE per token. This lesson reads the patterns, builds a uniform-vs-dynamic frame sampler, and evaluates on temporal grounding tasks.
|
||||
|
||||
**Type:** Build
|
||||
**Languages:** Python (stdlib, frame sampler + temporal-grounding evaluator)
|
||||
**Prerequisites:** Phase 12 · 08 (LLaVA-OneVision)
|
||||
**Time:** ~180 minutes
|
||||
|
||||
## Learning Objectives
|
||||
|
||||
- Explain why temporal positional encoding changes video VLM performance independently of the vision encoder.
|
||||
- Compare uniform, dynamic-FPS, and event-driven frame sampling on tokens-per-second vs grounding accuracy.
|
||||
- Describe Q-former-per-clip (Video-LLaMA) vs pooled-per-frame (Video-LLaVA) vs M-RoPE-per-token (Qwen2.5-VL) designs.
|
||||
- Name the four video benchmarks: VideoMME, TempCompass, EgoSchema, Video-MMMU.
|
||||
|
||||
## The Problem
|
||||
|
||||
A 1-minute video at 30 FPS is 1800 frames. At 196 visual tokens per frame (ViT-B at 224), that is 352k tokens — larger than any 2024-era LLM context.
|
||||
|
||||
Three reduction strategies exist:
|
||||
|
||||
1. Subsample frames (1-8 FPS depending on content).
|
||||
2. Pool each frame's patch tokens aggressively (3x3 or 4x4 bilinear pool).
|
||||
3. Compress via a Q-former that takes a 16-frame clip and outputs 64 tokens.
|
||||
|
||||
Each trade-off is different. Subsampling loses temporal detail. Pooling loses spatial detail. Q-former loses both a little but saves tokens.
|
||||
|
||||
Temporal position encoding is the other axis: how does the model know frame 5 came before frame 6? Options include simple 1D temporal RoPE (Video-LLaMA), learned temporal embeddings (Video-LLaVA), and TMRoPE (Qwen2.5-VL, full 3D).
|
||||
|
||||
## The Concept
|
||||
|
||||
### Video-LLaMA: Q-former per clip + audio branch
|
||||
|
||||
Video-LLaMA (2023) was the first open video-LLM. Architecture:
|
||||
|
||||
- 16-frame clips at 2 FPS (so 8 seconds).
|
||||
- Per-frame ViT features -> Video Q-former that cross-attends over all 16 frames -> 32 learned queries -> LLM.
|
||||
- Parallel audio branch: waveform -> ImageBind audio encoder -> Audio Q-former -> 32 queries -> LLM.
|
||||
|
||||
Strength: audio-visual joint reasoning. Weakness: fixed clip length, no arbitrary time grounding.
|
||||
|
||||
### VideoChat and Video-LLaVA
|
||||
|
||||
VideoChat kept the Video-LLaMA idea but dropped audio and simplified. Video-LLaVA (Lin et al., 2023) trained a single visual encoder on both images and video frames ("alignment before projection"), giving a unified representation. Both are frozen-CLIP-encoder + MLP + LLM.
|
||||
|
||||
Neither handles long video. Both are 8-16 frame systems.
|
||||
|
||||
### Qwen2.5-VL and TMRoPE
|
||||
|
||||
Qwen2.5-VL introduced TMRoPE — Temporal-Modality Rotary Position Embedding. Each patch token carries an (t, h, w) position where t is the actual timestamp (not frame index).
|
||||
|
||||
Key differences from simple temporal embedding:
|
||||
|
||||
- Absolute time, not index. The model sees "at 4.2 seconds" not "at frame 15."
|
||||
- Per-token rotation, not per-clip. Each visual token rotates independently by its timestamp.
|
||||
- Compatible with dynamic FPS. If you sample at 2 FPS here and 4 FPS there, TMRoPE handles the uneven spacing natively.
|
||||
|
||||
TMRoPE enables "at what second does the cat jump?" queries. The model can output "at 4.2 seconds." Video-LLaMA could only say "early in the clip."
|
||||
|
||||
### Frame sampling strategies
|
||||
|
||||
Uniform: sample N frames evenly over duration. Simple, loses motion peaks.
|
||||
|
||||
Dynamic FPS: sample adaptively based on motion intensity. Optical flow or frame differencing picks high-motion segments for denser sampling. Qwen2.5-VL trains on this.
|
||||
|
||||
Event-driven: run a lightweight detector, sample more where action happens. Used by VideoAgent.
|
||||
|
||||
Keyframe + context: sample at shot boundaries + a few adjacent frames. Used for cinematic content.
|
||||
|
||||
### Pooling per frame
|
||||
|
||||
At 1 FPS and 576 tokens per frame, a 5-minute clip is 172,800 tokens. Doable with Qwen2.5-VL-72B's 128k context but expensive.
|
||||
|
||||
3x3 bilinear pool reduces to 64 tokens per frame -> 19,200 tokens for 5 minutes. Sweet spot for most tasks.
|
||||
|
||||
Pool more aggressively (6x6 -> 16 tokens per frame) for agent workflows where spatial detail matters less.
|
||||
|
||||
### The four video benchmarks
|
||||
|
||||
- VideoMME: comprehensive video understanding, short + medium + long.
|
||||
- TempCompass: fine-grained temporal reasoning, "before" / "after" questions.
|
||||
- EgoSchema: long-horizon first-person video.
|
||||
- Video-MMMU: multimodal multi-discipline video questions.
|
||||
|
||||
A full video-VLM evaluation hits all four. They stress different axes — TempCompass is all about ordering, EgoSchema is about 3+ minute reasoning, VideoMME spans durations.
|
||||
|
||||
### Grounding output formats
|
||||
|
||||
Output formats for temporal grounding:
|
||||
|
||||
- Free text: "The cat jumps around the 4-second mark." Easy to parse but imprecise.
|
||||
- Structured JSON: `{"event": "jump", "start": 4.1, "end": 4.3}`. Qwen2.5-VL trains this.
|
||||
- Token-based: special `<time>4.1</time>` tokens interleaved with the answer. Qwen2.5-VL's internal format.
|
||||
|
||||
Token-based is most accurate for downstream use. Qwen2.5-VL's JSON output format parses directly.
|
||||
|
||||
### 2026 best practice
|
||||
|
||||
For video VLMs in 2026:
|
||||
|
||||
- Encoder: SigLIP 2 with M-RoPE or TMRoPE (Qwen2.5-VL).
|
||||
- Frame sampling: dynamic FPS (1-4 depending on motion) with max-frame cap.
|
||||
- Per-frame pooling: 3x3 bilinear.
|
||||
- Output: structured JSON with time + event fields.
|
||||
- Benchmarks: VideoMME + TempCompass for general; EgoSchema for long-horizon.
|
||||
|
||||
## Use It
|
||||
|
||||
`code/main.py` includes:
|
||||
|
||||
- Uniform and dynamic-FPS frame samplers.
|
||||
- A toy temporal-grounding evaluator: given a "ground truth" event at time T and a model output, score accuracy with tolerance.
|
||||
- A comparison across Video-LLaMA (16 frames, Q-former), Video-LLaVA (8 frames, MLP), Qwen2.5-VL (dynamic FPS + TMRoPE).
|
||||
|
||||
## Ship It
|
||||
|
||||
This lesson produces `outputs/skill-video-vlm-frame-planner.md`. Given a video task (monitoring, action recognition, temporal grounding, summarization), it picks the frame sampler, pooling factor, output format, and expected accuracy tier.
|
||||
|
||||
## Exercises
|
||||
|
||||
1. For a 3-minute cooking demo, pick uniform vs dynamic FPS. Justify with a token count.
|
||||
|
||||
2. TMRoPE adds what specifically that a simple temporal embedding table cannot do?
|
||||
|
||||
3. Write a JSON schema for temporal grounding that a VLM can learn to emit. Include error cases.
|
||||
|
||||
4. Read Video-LLaVA's Section 3 on "Alignment Before Projection." Why is this better than training separate image and video encoders?
|
||||
|
||||
5. Given the VideoMME leaderboard, what is the gap between the top open model and the top proprietary model as of 2026? How much of that gap is attributable to temporal encoding vs base LLM scale?
|
||||
|
||||
## Key Terms
|
||||
|
||||
| Term | What people say | What it actually means |
|
||||
|------|-----------------|------------------------|
|
||||
| Temporal grounding | "Time-localized answers" | VLM outputs a specific timestamp range for when an event happens |
|
||||
| TMRoPE | "Time-Multimodal RoPE" | 3D rotary position with absolute timestamps, used by Qwen2.5-VL |
|
||||
| Dynamic FPS | "Motion-aware sampling" | Sample more frames in high-motion segments, fewer in static ones |
|
||||
| Frame pooling | "Spatial compress per frame" | Reduce patches per frame with bilinear interpolation before the LLM |
|
||||
| Video Q-former | "Clip compressor" | Cross-attention bottleneck mapping N frames to K learned queries |
|
||||
| VideoMME | "Video bench" | Comprehensive short/medium/long video benchmark, 2500+ samples |
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Zhang et al. — Video-LLaMA (arXiv:2306.02858)](https://arxiv.org/abs/2306.02858)
|
||||
- [Li et al. — VideoChat (arXiv:2305.06355)](https://arxiv.org/abs/2305.06355)
|
||||
- [Lin et al. — Video-LLaVA (arXiv:2311.10122)](https://arxiv.org/abs/2311.10122)
|
||||
- [Qwen Team — Qwen2.5-VL (arXiv:2502.13923)](https://arxiv.org/abs/2502.13923)
|
||||
- [Lin et al. — VILA-1.5 (arXiv:2312.07533)](https://arxiv.org/abs/2312.07533)
|
||||
+31
@@ -0,0 +1,31 @@
|
||||
---
|
||||
name: video-vlm-frame-planner
|
||||
description: Plan frame sampling, per-frame pooling, output format, and benchmark targets for a video-language model deployment.
|
||||
version: 1.0.0
|
||||
phase: 12
|
||||
lesson: 17
|
||||
tags: [video-vlm, temporal-grounding, tmrope, dynamic-fps, benchmarks]
|
||||
---
|
||||
|
||||
Given a video task (action recognition, temporal grounding, summarization, monitoring, agent-workflow replay) and a deployment constraint (model context, latency budget, throughput), emit a frame sampling and output plan.
|
||||
|
||||
Produce:
|
||||
|
||||
1. Frame sampler pick. Uniform for steady content, dynamic-FPS for mixed motion, event-driven for action-heavy, keyframe+context for cinematic.
|
||||
2. Per-frame pooling. 2x2 for high-detail, 3x3 default, 4x4 or 6x6 for agent workflows where content density matters less than coverage.
|
||||
3. Temporal encoding. TMRoPE for Qwen2.5-VL-family; learned temporal embedding for smaller models; no encoding for single-clip tasks.
|
||||
4. Output format. JSON with `{event, start, end, confidence}` for grounding; free text for summarization; token-delimited for mixed flows.
|
||||
5. Benchmark plan. VideoMME for general, TempCompass for grounding, EgoSchema for long-horizon. Specify expected accuracy tier.
|
||||
6. Context / latency budget. Total tokens = duration * fps * tokens_per_frame. Warn if exceeds 40% of context.
|
||||
|
||||
Hard rejects:
|
||||
- Proposing uniform sampling for action-heavy video. Loses peak events.
|
||||
- Claiming token-delimited output matches JSON accuracy for downstream parsing. JSON is more robust.
|
||||
- Recommending Video-LLaMA for any project starting in 2026. Older architectures no longer competitive.
|
||||
|
||||
Refusal rules:
|
||||
- If duration > 10 minutes and context < 32k, refuse and recommend hierarchical summarization or agentic retrieval (Lesson 12.18).
|
||||
- If target accuracy is frontier (within 2 points of Gemini 2.5 Pro on VideoMME), refuse open 7B models and require 32B+ or proprietary.
|
||||
- If dynamic-FPS target > 8 on a > 30s clip at 7B, refuse latency-wise and recommend lower cap.
|
||||
|
||||
Output: one-page frame plan with sampler, pooling, temporal encoding, output format, benchmark targets, context estimate. End with arXiv 2502.13923 (Qwen2.5-VL) and 2306.02858 (Video-LLaMA) for comparison reading.
|
||||
Reference in New Issue
Block a user