Files
modeling_zhaocui/deep_learning/Q1

Q1: feature extraction and temporal alignment

Q1's required output is a reproducible set of three-modality features extracted from all 100 raw videos in Attachment 1, with traceable timestamps and at least one inspectable text/audio/video example. The four-method comparison and learned probes are optional research extensions; they follow the extraction and submission artifacts rather than replacing them.

For a nontechnical introduction to the extraction steps and their limits, see 方法说明_零基础版.md.

Environment

The project is managed with uv. The Linux environment selects PyTorch's CUDA 13.0 wheel; the checked machine is an RTX 5070 Ti under Fedora WSL.

cd Q1
uv sync
uv run python -c "import torch; print(torch.__version__, torch.cuda.is_available())"

uv.lock records the exact environment. The root data directory remains outside the Python environment and is not copied into the package.

Q1-E1: extract the 100 raw-video samples

cd Q1
uv run python -m q1.extract_features --output-dir outputs/q1_features

The first run downloads the pinned-by-name Hugging Face model revisions to the local cache and the MediaPipe face model to Q1/models/. Use --resume to continue an interrupted run. The supplied transcripts remain unchanged.

Each compressed NPZ keeps variable-length features without disk padding:

  • Text: 768-d BERT vectors, one mean-pooled vector per supplied whitespace word.
  • Audio: 25-d eGeMAPSv02 low-level descriptors from mono 16 kHz audio, with frame-center timestamps.
  • Vision: 192-d DeiT-Tiny CLS embeddings for every frame sampled at 5 fps; optional 52-d MediaPipe face blendshapes have a separate validity mask.
  • Alignment: CTC Viterbi word intervals in seconds from clip start, plus word-level mean Audio and Vision features. Missing CTC word spans are interpolated only as a flagged fallback (word_alignment_valid=false).

Outputs are written to outputs/q1_features/: one NPZ per original video_id/clip_id, JSONL sample logs, a 100-row sample table, a 300-row sample-by-modality table, a run manifest, and a typical-sample figure under reports/. Missing faces do not remove the general Vision embedding; missing modality values remain masked and are logged.

The extracted feature files and reports are approximately 10 MB before any optional Q1 method-comparison outputs. Keep the final combined Q1/Q2/Q3 attachments under the problem's 50 MB cap.

Q1-E0: sample coverage audit

cd Q1
uv run python -m q1.audit

The audit checks that the label workbook contains exactly 100 unique (video_id, clip_id) pairs, resolves each pair to one MP4, counts the 37 video_id groups, probes each MP4's duration and audio/video streams, and checks polarity labels against the sign of the continuous label. It writes outputs/audit/manifest.csv and outputs/audit/coverage_summary.json.

Required Q1 deliverables

  1. Three-modality features generated from the original 100 videos and their supplied transcripts. Keep every original row and label; if a stream needs a mask or truncation, retain the source and record the rule.
  2. Per-sample and per-modality records for effective duration, feature dimension, valid length, padding rule, alignment granularity, and mapping back to source time.
  3. A 100-sample results table and a typical-sample view connecting transcript text, aligned speech intervals, video frame times, and feature positions.
  4. A run manifest with tool/model versions, parameters, logs, and commands.
  5. Compressed feature files that fit the competition's 50 MB total-attachment limit together with the Q2/Q3 code and result files.

The Attachment 1 media and labels are fixed input. The labels are for the required inventory and later prediction probes; they are not a required training target for the primary Q1 alignment model.

Common feature and alignment contract

Each modality is supplied as a padded SequenceBatch:

  • features: [B, L_m, D_m] feature values;
  • times: [B, L_m] timestamps in seconds from clip start;
  • valid: [B, L_m] boolean mask; padded positions are false.

The duration vector is [B]. Every method returns AlignmentOutput with weights[m] shaped [B, K, L_m]. Rows sum to one and columns for padding stay zero. aligned[m] is the weighted or learned representation on the shared K=50 grid. M1 takes one ordered [word_count, 2] forced-word-interval array per sample; its 50 grid intervals partition transcript word order and preserve the measured word times. M2 uses equal-duration windows.

M3 uses text-order slots as queries for Audio and Vision. M4 learns K shared latent queries and attends separately to all three sequences. Both use the same module dimensions and can be trained with alignment_training_loss; M1/M2 have no training loss.

Initial metrics

q1.metrics provides normalized expected-time trajectories, monotonicity violation rate, normalized attention entropy, and shortest contiguous 80% attention width. Entropy and width are diagnostics, not one-direction ranking scores. A flat trajectory has zero backward violations too, so inspect its time span as well; otherwise a collapsed alignment can look monotone. Grid-slot retrieval is also available as a representation-consistency probe; because its positive is defined by the shared grid index, it is not independent evidence of temporal correctness. Use a human-labeled event subset for direct temporal claims.

Optional method-comparison rules

  • Extract all three modalities once with the same extractor versions and keep source timestamps, model IDs, and preprocessing parameters in the manifest.
  • Split clips by video_id; never allow clips from one source video into both training and validation folds.
  • Fit learned M3/M4 models only on training folds. Freeze them before probes.
  • Keep reconstruction and emotion probes identical across alignment methods.
  • Do not use emotion labels in the primary alignment loss.
  • Record all seeds, fold IDs, attention matrices, masks, and metrics alongside each output.

M1–M4 grouped comparison runner

Run the optional method comparison after feature extraction and the coverage audit are complete:

cd Q1
uv run python -m q1.compare_methods

The runner selects CUDA when available, fits feature normalization on each training fold only, and uses five grouped folds by video_id. M3/M4 share the same masked-reconstruction, cross-modal contrastive, and temporal-monotonicity objective; emotion labels are withheld from alignment training. It writes fold checkpoints, out-of-fold alignment matrices, trajectories, retrieval and reconstruction probes, and a frozen emotion probe under outputs/method_comparison/. Retrieval uses same-grid positives, so it is a representation-consistency check rather than independent temporal ground truth. Direct human IoU/MATE scores require a separately annotated event set. On Fedora WSL, PyTorch's Triton CUDA backward path also needs a system C compiler and Python headers (sudo dnf install gcc python3-devel); Python packages remain managed by uv.

M3/M4 attention-collapse follow-up

After the first comparison showed nearly uniform M3/M4 attention, the staged follow-up changes only learned alignment losses and probes. It reuses the saved feature files and original video-grouped folds, and leaves M1/M2 untouched:

cd Q1
uv run python -m q1.train_alignment_variants

The runner executes v2-a (Span), v2-b (Span + far-slot Diversity), and v2-c (Span + Diversity + weak Temporal Band) with the same folds and three seeds. It records C_row during validation, tests within-clip temporal retrieval, and compares frozen masked reconstruction with shuffled cross-modal slots. Full checkpoints and per-sample alignment matrices are written under outputs/alignment_v2/; use outputs/alignment_v2/report_bundle/ for a lightweight shareable summary. RoPE/Gaussian-bias experiments should wait until the attention maps show time-progressing local bands.

M3/M4 learnability diagnostics

Before running more architecture or positional-encoding sweeps, diagnose the attention-collapse behavior with the small D0–D5 suite plus the D1-S synthetic control:

cd Q1
uv run python -m q1.alignment_debug
uv run python -m q1.synthetic_alignment_sanity
uv run python -m q1.alignment_key_position_debug
uv run python -m q1.alignment_heldout_debug
uv run python -m q1.summarize_alignment_debug

D0/D1 overfit one representative clip. D2/D3 use all 100 clips for an in-sample optimization diagnostic, with one seed. D1 compares M4 with and without fixed sinusoidal positions; D2 uses a timestamp-derived Gaussian attention target; D3 adds the same reconstruction and contrastive components used in the prior learned methods. The Gaussian target is only a weak temporal prior, not human alignment ground truth. The run records component losses and gradients for attention Q/K and M4 latent slots, then saves per-sample metrics, heatmaps, trajectories, and checkpoints under outputs/alignment_debug/. The optional synthetic control uses only [t, t², 1] source features to verify that M3/M4 can fit a known time band when position information is explicit. The source-time diagnostic adds Fourier time identity to attention keys while keeping the projected content as values; the held-out runner uses grouped fold 1 by default and fits feature normalization on its training videos only. Its --fold and --output-dir options allow the same diagnostic to run on the remaining grouped folds. Summarize the existing run without retraining with the second command; it also builds a compact outputs/alignment_debug/report_bundle/ without checkpoints. The diagnostics found nonzero gradients. D4/D5 show that M4 Audio/Vision can form local bands with explicit time identity, while M3 and the M4 Text branch still need work. Keep follow-up focused on those branches instead of broad positional-encoding sweeps; see RESULTS.md for metrics and interpretation.

Same-slot vs shifted-time correspondence evaluation

Temporal attention bands only show where a slot looks. To check whether the frozen content in the slot can match the other modalities, first create D5 checkpoints for all five video_id-grouped folds with q1.alignment_heldout_debug, then run:

cd Q1
uv run python -m q1.correspondence_eval --device cuda --epochs 40 --seed 42

The evaluator applies the same train-only linear probe to M1, M2, and both time-code settings of M3/M4. It reports within-video shifted-similarity curves, same-video temporal retrieval, and matched-vs-shifted AUC on held-out videos. The probe learns same-slot matching from training clips only; its held-out scores indicate whether that matchability transfers, not whether semantic alignment has independent ground truth. Outputs and grouped confidence intervals are written to outputs/correspondence_eval/; see RESULTS.md for the current results.

M4 Shared Latent structural and functional re-evaluation

Once the five grouped D5 M4_sourceTime checkpoints exist, re-evaluate M4 without retraining its alignment model:

cd Q1
uv run python -m q1.m4_shared_latent_eval --device cuda --seed 42 --probe-epochs 40 --decoder-epochs 40 --shuffle-repeats 20

The evaluator checks self-Gram structure, derives six modality-to-modality maps through the shared latent slots, measures physical-time trajectories, cycle and triangle consistency, and runs held-out content-only correspondence, within-clip retrieval, independent content-shuffle, and aligned-versus-shifted reconstruction probes. It fits only these small evaluation probes on training video groups; no emotion labels enter them. The reconstruction target is the M4 attention-pooled content representation, not raw audio or video. Because D5 used a timestamp-derived Gaussian prior, its attention weights may carry time information indirectly even though the probe receives no explicit time code or slot ID. Time structure and temporal maps therefore are not independent semantic ground truth. Human event IoU/center error still requires annotated events. CSVs, checkpoints, the run manifest, and figures are written to outputs/m4_shared_latent_eval/; see RESULTS.md for the five-fold results and interpretation.

TSFA coarse-to-fine correspondence experiment

After the five grouped D5 checkpoints exist, freeze M4's source-time branch and train a content-only local attention branch. The main variant searches within ±0.10 normalized video time around each frozen M4 temporal center. A multiply variant applies the M4 attention weights as an additional prior; random-window and global-window controls test whether the temporal candidate range matters. The original BERT text, eGeMAPS audio, and DeiT image features stay fixed. Emotion labels are not used.

cd Q1
uv run python -m q1.tsfa_experiment --device cuda --seed 42
uv run python -m q1.tsfa_alignment_only_eval --device cuda --seed 42

The first command fits one semantic branch and equal-seed correspondence probes per fold, then evaluates all seven methods on held-out video_id groups. The second pools the same normalized raw source features with each method's alignment matrix before fitting equal probes. It separates source-position selection from native model value/output projections. Both commands cover all 100 clips exactly once as held-out examples. Summary CSVs, per-clip metrics, paired video-group confidence intervals, monotonicity/span/entropy diagnostics, plots, manifests, and checkpoints are written to outputs/tsfa/; the small report_bundle/ contains shareable summaries and figures. See RESULTS.md for the measured tradeoff and interpretation limits.