287 lines
14 KiB
Markdown
287 lines
14 KiB
Markdown
# Q1: feature extraction and temporal alignment
|
||
|
||
Q1's required output is a reproducible set of three-modality features extracted
|
||
from all 100 raw videos in Attachment 1, with traceable timestamps and at least
|
||
one inspectable text/audio/video example. The four-method comparison and
|
||
learned probes are optional research extensions; they follow the extraction and
|
||
submission artifacts rather than replacing them.
|
||
|
||
For a nontechnical introduction to the extraction steps and their limits, see [方法说明_零基础版.md](方法说明_零基础版.md).
|
||
|
||
## Environment
|
||
|
||
The project is managed with `uv`. The Linux environment selects PyTorch's CUDA
|
||
13.0 wheel; the checked machine is an RTX 5070 Ti under Fedora WSL.
|
||
|
||
```bash
|
||
cd Q1
|
||
uv sync
|
||
uv run python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
|
||
```
|
||
|
||
`uv.lock` records the exact environment. The root data directory remains
|
||
outside the Python environment and is not copied into the package.
|
||
|
||
## Q1-E1: extract the 100 raw-video samples
|
||
|
||
```bash
|
||
cd Q1
|
||
uv run python -m q1.extract_features --output-dir outputs/q1_features
|
||
```
|
||
|
||
The first run downloads the pinned-by-name Hugging Face model revisions to the
|
||
local cache and the MediaPipe face model to `Q1/models/`. Use `--resume` to
|
||
continue an interrupted run. The supplied transcripts remain unchanged.
|
||
|
||
Each compressed NPZ keeps variable-length features without disk padding:
|
||
|
||
- Text: 768-d BERT vectors, one mean-pooled vector per supplied whitespace word.
|
||
- Audio: 25-d eGeMAPSv02 low-level descriptors from mono 16 kHz audio, with
|
||
frame-center timestamps.
|
||
- Vision: 192-d DeiT-Tiny CLS embeddings for every frame sampled at 5 fps;
|
||
optional 52-d MediaPipe face blendshapes have a separate validity mask.
|
||
- Alignment: CTC Viterbi word intervals in seconds from clip start, plus
|
||
word-level mean Audio and Vision features. Missing CTC word spans are
|
||
interpolated only as a flagged fallback (`word_alignment_valid=false`).
|
||
|
||
Outputs are written to `outputs/q1_features/`: one NPZ per original
|
||
`video_id/clip_id`, JSONL sample logs, a 100-row sample table, a 300-row
|
||
sample-by-modality table, a run manifest, and a typical-sample figure under
|
||
`reports/`. Missing faces do not remove the general Vision embedding; missing
|
||
modality values remain masked and are logged.
|
||
|
||
The extracted feature files and reports are approximately 10 MB before any
|
||
optional Q1 method-comparison outputs. Keep the final combined Q1/Q2/Q3
|
||
attachments under the problem's 50 MB cap.
|
||
|
||
## Q1-E0: sample coverage audit
|
||
|
||
```bash
|
||
cd Q1
|
||
uv run python -m q1.audit
|
||
```
|
||
|
||
The audit checks that the label workbook contains exactly 100 unique
|
||
`(video_id, clip_id)` pairs, resolves each pair to one MP4, counts the 37
|
||
`video_id` groups, probes each MP4's duration and audio/video streams, and checks
|
||
polarity labels against the sign of the continuous label. It writes
|
||
`outputs/audit/manifest.csv` and `outputs/audit/coverage_summary.json`.
|
||
|
||
## Required Q1 deliverables
|
||
|
||
1. Three-modality features generated from the original 100 videos and their
|
||
supplied transcripts. Keep every original row and label; if a stream needs a
|
||
mask or truncation, retain the source and record the rule.
|
||
2. Per-sample and per-modality records for effective duration, feature
|
||
dimension, valid length, padding rule, alignment granularity, and mapping
|
||
back to source time.
|
||
3. A 100-sample results table and a typical-sample view connecting transcript
|
||
text, aligned speech intervals, video frame times, and feature positions.
|
||
4. A run manifest with tool/model versions, parameters, logs, and commands.
|
||
5. Compressed feature files that fit the competition's 50 MB total-attachment
|
||
limit together with the Q2/Q3 code and result files.
|
||
|
||
The Attachment 1 media and labels are fixed input. The labels are for the
|
||
required inventory and later prediction probes; they are not a required training
|
||
target for the primary Q1 alignment model.
|
||
|
||
## Common feature and alignment contract
|
||
|
||
Each modality is supplied as a padded `SequenceBatch`:
|
||
|
||
- `features`: `[B, L_m, D_m]` feature values;
|
||
- `times`: `[B, L_m]` timestamps in seconds from clip start;
|
||
- `valid`: `[B, L_m]` boolean mask; padded positions are false.
|
||
|
||
The duration vector is `[B]`. Every method returns `AlignmentOutput` with
|
||
`weights[m]` shaped `[B, K, L_m]`. Rows sum to one and columns for padding stay
|
||
zero. `aligned[m]` is the weighted or learned representation on the shared
|
||
`K=50` grid. M1 takes one ordered `[word_count, 2]` forced-word-interval array
|
||
per sample; its 50 grid intervals partition transcript word order and preserve
|
||
the measured word times. M2 uses equal-duration windows.
|
||
|
||
M3 uses text-order slots as queries for Audio and Vision. M4 learns K shared
|
||
latent queries and attends separately to all three sequences. Both use the same
|
||
module dimensions and can be trained with `alignment_training_loss`; M1/M2 have
|
||
no training loss.
|
||
|
||
## Initial metrics
|
||
|
||
`q1.metrics` provides normalized expected-time trajectories, monotonicity
|
||
violation rate, normalized attention entropy, and shortest contiguous 80%
|
||
attention width. Entropy and width are diagnostics, not one-direction ranking
|
||
scores. A flat trajectory has zero backward violations too, so inspect its
|
||
time span as well; otherwise a collapsed alignment can look monotone. Grid-slot
|
||
retrieval is also available as a representation-consistency probe; because its
|
||
positive is defined by the shared grid index, it is not independent evidence
|
||
of temporal correctness. Use a human-labeled event subset for direct temporal
|
||
claims.
|
||
|
||
## Optional method-comparison rules
|
||
|
||
- Extract all three modalities once with the same extractor versions and keep
|
||
source timestamps, model IDs, and preprocessing parameters in the manifest.
|
||
- Split clips by `video_id`; never allow clips from one source video into both
|
||
training and validation folds.
|
||
- Fit learned M3/M4 models only on training folds. Freeze them before probes.
|
||
- Keep reconstruction and emotion probes identical across alignment methods.
|
||
- Do not use emotion labels in the primary alignment loss.
|
||
- Record all seeds, fold IDs, attention matrices, masks, and metrics alongside
|
||
each output.
|
||
|
||
## M1–M4 grouped comparison runner
|
||
|
||
Run the optional method comparison after feature extraction and the coverage
|
||
audit are complete:
|
||
|
||
```bash
|
||
cd Q1
|
||
uv run python -m q1.compare_methods
|
||
```
|
||
|
||
The runner selects CUDA when available, fits feature normalization on each
|
||
training fold only, and uses five grouped folds by `video_id`. M3/M4 share the
|
||
same masked-reconstruction, cross-modal contrastive, and temporal-monotonicity
|
||
objective; emotion labels are withheld from alignment training. It writes
|
||
fold checkpoints, out-of-fold alignment matrices, trajectories, retrieval and
|
||
reconstruction probes, and a frozen emotion probe under
|
||
`outputs/method_comparison/`. Retrieval uses same-grid positives, so it is a
|
||
representation-consistency check rather than independent temporal ground
|
||
truth. Direct human IoU/MATE scores require a separately annotated event set.
|
||
On Fedora WSL, PyTorch's Triton CUDA backward path also needs a system C
|
||
compiler and Python headers (`sudo dnf install gcc python3-devel`); Python
|
||
packages remain managed by `uv`.
|
||
|
||
## M3/M4 attention-collapse follow-up
|
||
|
||
After the first comparison showed nearly uniform M3/M4 attention, the staged
|
||
follow-up changes only learned alignment losses and probes. It reuses the saved
|
||
feature files and original video-grouped folds, and leaves M1/M2 untouched:
|
||
|
||
```bash
|
||
cd Q1
|
||
uv run python -m q1.train_alignment_variants
|
||
```
|
||
|
||
The runner executes v2-a (Span), v2-b (Span + far-slot Diversity), and v2-c
|
||
(Span + Diversity + weak Temporal Band) with the same folds and three seeds.
|
||
It records `C_row` during validation, tests within-clip temporal retrieval,
|
||
and compares frozen masked reconstruction with shuffled cross-modal slots.
|
||
Full checkpoints and per-sample alignment matrices are written under
|
||
`outputs/alignment_v2/`; use `outputs/alignment_v2/report_bundle/` for a
|
||
lightweight shareable summary. RoPE/Gaussian-bias experiments should wait until
|
||
the attention maps show time-progressing local bands.
|
||
|
||
## M3/M4 learnability diagnostics
|
||
|
||
Before running more architecture or positional-encoding sweeps, diagnose the
|
||
attention-collapse behavior with the small D0–D5 suite plus the D1-S synthetic
|
||
control:
|
||
|
||
```bash
|
||
cd Q1
|
||
uv run python -m q1.alignment_debug
|
||
uv run python -m q1.synthetic_alignment_sanity
|
||
uv run python -m q1.alignment_key_position_debug
|
||
uv run python -m q1.alignment_heldout_debug
|
||
uv run python -m q1.summarize_alignment_debug
|
||
```
|
||
|
||
D0/D1 overfit one representative clip. D2/D3 use all 100 clips for an
|
||
in-sample optimization diagnostic, with one seed. D1 compares M4 with and
|
||
without fixed sinusoidal positions; D2 uses a timestamp-derived Gaussian
|
||
attention target; D3 adds the same reconstruction and contrastive components
|
||
used in the prior learned methods. The Gaussian target is only a weak temporal
|
||
prior, not human alignment ground truth. The run records component losses and
|
||
gradients for attention Q/K and M4 latent slots, then saves per-sample metrics,
|
||
heatmaps, trajectories, and checkpoints under `outputs/alignment_debug/`.
|
||
The optional synthetic control uses only `[t, t², 1]` source features to verify
|
||
that M3/M4 can fit a known time band when position information is explicit.
|
||
The source-time diagnostic adds Fourier time identity to attention keys while
|
||
keeping the projected content as values; the held-out runner uses grouped fold
|
||
1 by default and fits feature normalization on its training videos only. Its
|
||
`--fold` and `--output-dir` options allow the same diagnostic to run on the
|
||
remaining grouped folds.
|
||
Summarize the existing run without retraining with the second command; it also
|
||
builds a compact `outputs/alignment_debug/report_bundle/` without checkpoints.
|
||
The diagnostics found nonzero gradients. D4/D5 show that M4 Audio/Vision can
|
||
form local bands with explicit time identity, while M3 and the M4 Text branch
|
||
still need work. Keep follow-up focused on those branches instead of broad
|
||
positional-encoding sweeps; see [RESULTS.md](RESULTS.md) for metrics and
|
||
interpretation.
|
||
|
||
## Same-slot vs shifted-time correspondence evaluation
|
||
|
||
Temporal attention bands only show where a slot looks. To check whether the
|
||
frozen content in the slot can match the other modalities, first create D5
|
||
checkpoints for all five `video_id`-grouped folds with
|
||
`q1.alignment_heldout_debug`, then run:
|
||
|
||
```bash
|
||
cd Q1
|
||
uv run python -m q1.correspondence_eval --device cuda --epochs 40 --seed 42
|
||
```
|
||
|
||
The evaluator applies the same train-only linear probe to M1, M2, and both
|
||
time-code settings of M3/M4. It reports within-video shifted-similarity curves,
|
||
same-video temporal retrieval, and matched-vs-shifted AUC on held-out videos.
|
||
The probe learns same-slot matching from training clips only; its held-out
|
||
scores indicate whether that matchability transfers, not whether semantic
|
||
alignment has independent ground truth. Outputs and grouped confidence
|
||
intervals are written to `outputs/correspondence_eval/`; see
|
||
[RESULTS.md](RESULTS.md) for the current results.
|
||
|
||
## M4 Shared Latent structural and functional re-evaluation
|
||
|
||
Once the five grouped D5 `M4_sourceTime` checkpoints exist, re-evaluate M4
|
||
without retraining its alignment model:
|
||
|
||
```bash
|
||
cd Q1
|
||
uv run python -m q1.m4_shared_latent_eval --device cuda --seed 42 --probe-epochs 40 --decoder-epochs 40 --shuffle-repeats 20
|
||
```
|
||
|
||
The evaluator checks self-Gram structure, derives six modality-to-modality
|
||
maps through the shared latent slots, measures physical-time trajectories,
|
||
cycle and triangle consistency, and runs held-out content-only correspondence,
|
||
within-clip retrieval, independent content-shuffle, and aligned-versus-shifted
|
||
reconstruction probes. It fits only these small evaluation probes on training
|
||
video groups; no emotion labels enter them. The reconstruction target is the
|
||
M4 attention-pooled content representation, not raw audio or video. Because
|
||
D5 used a timestamp-derived Gaussian prior, its attention weights may carry
|
||
time information indirectly even though the probe receives no explicit time
|
||
code or slot ID. Time structure and temporal maps therefore are not
|
||
independent semantic ground truth. Human event IoU/center error still requires
|
||
annotated events. CSVs, checkpoints, the run manifest, and figures are written
|
||
to `outputs/m4_shared_latent_eval/`; see [RESULTS.md](RESULTS.md) for the
|
||
five-fold results and interpretation.
|
||
|
||
## TSFA coarse-to-fine correspondence experiment
|
||
|
||
After the five grouped D5 checkpoints exist, freeze M4's source-time branch and
|
||
train a content-only local attention branch. The main variant searches within
|
||
`±0.10` normalized video time around each frozen M4 temporal center. A
|
||
multiply variant applies the M4 attention weights as an additional prior;
|
||
random-window and global-window controls test whether the temporal candidate
|
||
range matters. The original BERT text, eGeMAPS audio, and DeiT image features
|
||
stay fixed. Emotion labels are not used.
|
||
|
||
```bash
|
||
cd Q1
|
||
uv run python -m q1.tsfa_experiment --device cuda --seed 42
|
||
uv run python -m q1.tsfa_alignment_only_eval --device cuda --seed 42
|
||
```
|
||
|
||
The first command fits one semantic branch and equal-seed correspondence probes
|
||
per fold, then evaluates all seven methods on held-out `video_id` groups. The
|
||
second pools the *same normalized raw source features* with each method's
|
||
alignment matrix before fitting equal probes. It separates source-position
|
||
selection from native model value/output projections. Both commands cover all
|
||
100 clips exactly once as held-out examples. Summary CSVs, per-clip metrics,
|
||
paired video-group confidence intervals, monotonicity/span/entropy diagnostics,
|
||
plots, manifests, and checkpoints are written to `outputs/tsfa/`; the small
|
||
`report_bundle/` contains shareable
|
||
summaries and figures. See [RESULTS.md](RESULTS.md) for the measured tradeoff
|
||
and interpretation limits.
|
||
|