14 KiB
Q1: feature extraction and temporal alignment
Q1's required output is a reproducible set of three-modality features extracted from all 100 raw videos in Attachment 1, with traceable timestamps and at least one inspectable text/audio/video example. The four-method comparison and learned probes are optional research extensions; they follow the extraction and submission artifacts rather than replacing them.
For a nontechnical introduction to the extraction steps and their limits, see 方法说明_零基础版.md.
Environment
The project is managed with uv. The Linux environment selects PyTorch's CUDA
13.0 wheel; the checked machine is an RTX 5070 Ti under Fedora WSL.
cd Q1
uv sync
uv run python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
uv.lock records the exact environment. The root data directory remains
outside the Python environment and is not copied into the package.
Q1-E1: extract the 100 raw-video samples
cd Q1
uv run python -m q1.extract_features --output-dir outputs/q1_features
The first run downloads the pinned-by-name Hugging Face model revisions to the
local cache and the MediaPipe face model to Q1/models/. Use --resume to
continue an interrupted run. The supplied transcripts remain unchanged.
Each compressed NPZ keeps variable-length features without disk padding:
- Text: 768-d BERT vectors, one mean-pooled vector per supplied whitespace word.
- Audio: 25-d eGeMAPSv02 low-level descriptors from mono 16 kHz audio, with frame-center timestamps.
- Vision: 192-d DeiT-Tiny CLS embeddings for every frame sampled at 5 fps; optional 52-d MediaPipe face blendshapes have a separate validity mask.
- Alignment: CTC Viterbi word intervals in seconds from clip start, plus
word-level mean Audio and Vision features. Missing CTC word spans are
interpolated only as a flagged fallback (
word_alignment_valid=false).
Outputs are written to outputs/q1_features/: one NPZ per original
video_id/clip_id, JSONL sample logs, a 100-row sample table, a 300-row
sample-by-modality table, a run manifest, and a typical-sample figure under
reports/. Missing faces do not remove the general Vision embedding; missing
modality values remain masked and are logged.
The extracted feature files and reports are approximately 10 MB before any optional Q1 method-comparison outputs. Keep the final combined Q1/Q2/Q3 attachments under the problem's 50 MB cap.
Q1-E0: sample coverage audit
cd Q1
uv run python -m q1.audit
The audit checks that the label workbook contains exactly 100 unique
(video_id, clip_id) pairs, resolves each pair to one MP4, counts the 37
video_id groups, probes each MP4's duration and audio/video streams, and checks
polarity labels against the sign of the continuous label. It writes
outputs/audit/manifest.csv and outputs/audit/coverage_summary.json.
Required Q1 deliverables
- Three-modality features generated from the original 100 videos and their supplied transcripts. Keep every original row and label; if a stream needs a mask or truncation, retain the source and record the rule.
- Per-sample and per-modality records for effective duration, feature dimension, valid length, padding rule, alignment granularity, and mapping back to source time.
- A 100-sample results table and a typical-sample view connecting transcript text, aligned speech intervals, video frame times, and feature positions.
- A run manifest with tool/model versions, parameters, logs, and commands.
- Compressed feature files that fit the competition's 50 MB total-attachment limit together with the Q2/Q3 code and result files.
The Attachment 1 media and labels are fixed input. The labels are for the required inventory and later prediction probes; they are not a required training target for the primary Q1 alignment model.
Common feature and alignment contract
Each modality is supplied as a padded SequenceBatch:
features:[B, L_m, D_m]feature values;times:[B, L_m]timestamps in seconds from clip start;valid:[B, L_m]boolean mask; padded positions are false.
The duration vector is [B]. Every method returns AlignmentOutput with
weights[m] shaped [B, K, L_m]. Rows sum to one and columns for padding stay
zero. aligned[m] is the weighted or learned representation on the shared
K=50 grid. M1 takes one ordered [word_count, 2] forced-word-interval array
per sample; its 50 grid intervals partition transcript word order and preserve
the measured word times. M2 uses equal-duration windows.
M3 uses text-order slots as queries for Audio and Vision. M4 learns K shared
latent queries and attends separately to all three sequences. Both use the same
module dimensions and can be trained with alignment_training_loss; M1/M2 have
no training loss.
Initial metrics
q1.metrics provides normalized expected-time trajectories, monotonicity
violation rate, normalized attention entropy, and shortest contiguous 80%
attention width. Entropy and width are diagnostics, not one-direction ranking
scores. A flat trajectory has zero backward violations too, so inspect its
time span as well; otherwise a collapsed alignment can look monotone. Grid-slot
retrieval is also available as a representation-consistency probe; because its
positive is defined by the shared grid index, it is not independent evidence
of temporal correctness. Use a human-labeled event subset for direct temporal
claims.
Optional method-comparison rules
- Extract all three modalities once with the same extractor versions and keep source timestamps, model IDs, and preprocessing parameters in the manifest.
- Split clips by
video_id; never allow clips from one source video into both training and validation folds. - Fit learned M3/M4 models only on training folds. Freeze them before probes.
- Keep reconstruction and emotion probes identical across alignment methods.
- Do not use emotion labels in the primary alignment loss.
- Record all seeds, fold IDs, attention matrices, masks, and metrics alongside each output.
M1–M4 grouped comparison runner
Run the optional method comparison after feature extraction and the coverage audit are complete:
cd Q1
uv run python -m q1.compare_methods
The runner selects CUDA when available, fits feature normalization on each
training fold only, and uses five grouped folds by video_id. M3/M4 share the
same masked-reconstruction, cross-modal contrastive, and temporal-monotonicity
objective; emotion labels are withheld from alignment training. It writes
fold checkpoints, out-of-fold alignment matrices, trajectories, retrieval and
reconstruction probes, and a frozen emotion probe under
outputs/method_comparison/. Retrieval uses same-grid positives, so it is a
representation-consistency check rather than independent temporal ground
truth. Direct human IoU/MATE scores require a separately annotated event set.
On Fedora WSL, PyTorch's Triton CUDA backward path also needs a system C
compiler and Python headers (sudo dnf install gcc python3-devel); Python
packages remain managed by uv.
M3/M4 attention-collapse follow-up
After the first comparison showed nearly uniform M3/M4 attention, the staged follow-up changes only learned alignment losses and probes. It reuses the saved feature files and original video-grouped folds, and leaves M1/M2 untouched:
cd Q1
uv run python -m q1.train_alignment_variants
The runner executes v2-a (Span), v2-b (Span + far-slot Diversity), and v2-c
(Span + Diversity + weak Temporal Band) with the same folds and three seeds.
It records C_row during validation, tests within-clip temporal retrieval,
and compares frozen masked reconstruction with shuffled cross-modal slots.
Full checkpoints and per-sample alignment matrices are written under
outputs/alignment_v2/; use outputs/alignment_v2/report_bundle/ for a
lightweight shareable summary. RoPE/Gaussian-bias experiments should wait until
the attention maps show time-progressing local bands.
M3/M4 learnability diagnostics
Before running more architecture or positional-encoding sweeps, diagnose the attention-collapse behavior with the small D0–D5 suite plus the D1-S synthetic control:
cd Q1
uv run python -m q1.alignment_debug
uv run python -m q1.synthetic_alignment_sanity
uv run python -m q1.alignment_key_position_debug
uv run python -m q1.alignment_heldout_debug
uv run python -m q1.summarize_alignment_debug
D0/D1 overfit one representative clip. D2/D3 use all 100 clips for an
in-sample optimization diagnostic, with one seed. D1 compares M4 with and
without fixed sinusoidal positions; D2 uses a timestamp-derived Gaussian
attention target; D3 adds the same reconstruction and contrastive components
used in the prior learned methods. The Gaussian target is only a weak temporal
prior, not human alignment ground truth. The run records component losses and
gradients for attention Q/K and M4 latent slots, then saves per-sample metrics,
heatmaps, trajectories, and checkpoints under outputs/alignment_debug/.
The optional synthetic control uses only [t, t², 1] source features to verify
that M3/M4 can fit a known time band when position information is explicit.
The source-time diagnostic adds Fourier time identity to attention keys while
keeping the projected content as values; the held-out runner uses grouped fold
1 by default and fits feature normalization on its training videos only. Its
--fold and --output-dir options allow the same diagnostic to run on the
remaining grouped folds.
Summarize the existing run without retraining with the second command; it also
builds a compact outputs/alignment_debug/report_bundle/ without checkpoints.
The diagnostics found nonzero gradients. D4/D5 show that M4 Audio/Vision can
form local bands with explicit time identity, while M3 and the M4 Text branch
still need work. Keep follow-up focused on those branches instead of broad
positional-encoding sweeps; see RESULTS.md for metrics and
interpretation.
Same-slot vs shifted-time correspondence evaluation
Temporal attention bands only show where a slot looks. To check whether the
frozen content in the slot can match the other modalities, first create D5
checkpoints for all five video_id-grouped folds with
q1.alignment_heldout_debug, then run:
cd Q1
uv run python -m q1.correspondence_eval --device cuda --epochs 40 --seed 42
The evaluator applies the same train-only linear probe to M1, M2, and both
time-code settings of M3/M4. It reports within-video shifted-similarity curves,
same-video temporal retrieval, and matched-vs-shifted AUC on held-out videos.
The probe learns same-slot matching from training clips only; its held-out
scores indicate whether that matchability transfers, not whether semantic
alignment has independent ground truth. Outputs and grouped confidence
intervals are written to outputs/correspondence_eval/; see
RESULTS.md for the current results.
M4 Shared Latent structural and functional re-evaluation
Once the five grouped D5 M4_sourceTime checkpoints exist, re-evaluate M4
without retraining its alignment model:
cd Q1
uv run python -m q1.m4_shared_latent_eval --device cuda --seed 42 --probe-epochs 40 --decoder-epochs 40 --shuffle-repeats 20
The evaluator checks self-Gram structure, derives six modality-to-modality
maps through the shared latent slots, measures physical-time trajectories,
cycle and triangle consistency, and runs held-out content-only correspondence,
within-clip retrieval, independent content-shuffle, and aligned-versus-shifted
reconstruction probes. It fits only these small evaluation probes on training
video groups; no emotion labels enter them. The reconstruction target is the
M4 attention-pooled content representation, not raw audio or video. Because
D5 used a timestamp-derived Gaussian prior, its attention weights may carry
time information indirectly even though the probe receives no explicit time
code or slot ID. Time structure and temporal maps therefore are not
independent semantic ground truth. Human event IoU/center error still requires
annotated events. CSVs, checkpoints, the run manifest, and figures are written
to outputs/m4_shared_latent_eval/; see RESULTS.md for the
five-fold results and interpretation.
TSFA coarse-to-fine correspondence experiment
After the five grouped D5 checkpoints exist, freeze M4's source-time branch and
train a content-only local attention branch. The main variant searches within
±0.10 normalized video time around each frozen M4 temporal center. A
multiply variant applies the M4 attention weights as an additional prior;
random-window and global-window controls test whether the temporal candidate
range matters. The original BERT text, eGeMAPS audio, and DeiT image features
stay fixed. Emotion labels are not used.
cd Q1
uv run python -m q1.tsfa_experiment --device cuda --seed 42
uv run python -m q1.tsfa_alignment_only_eval --device cuda --seed 42
The first command fits one semantic branch and equal-seed correspondence probes
per fold, then evaluates all seven methods on held-out video_id groups. The
second pools the same normalized raw source features with each method's
alignment matrix before fitting equal probes. It separates source-position
selection from native model value/output projections. Both commands cover all
100 clips exactly once as held-out examples. Summary CSVs, per-clip metrics,
paired video-group confidence intervals, monotonicity/span/entropy diagnostics,
plots, manifests, and checkpoints are written to outputs/tsfa/; the small
report_bundle/ contains shareable
summaries and figures. See RESULTS.md for the measured tradeoff
and interpretation limits.