Files

384 lines
18 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Q1: feature extraction and temporal alignment
Q1's required output is a reproducible set of three-modality features extracted
from all 100 raw videos in Attachment 1, with traceable timestamps and at least
one inspectable text/audio/video example. The four-method comparison and
learned probes are optional research extensions; they follow the extraction and
submission artifacts rather than replacing them.
For a nontechnical introduction to the extraction steps and their limits, see [方法说明_零基础版.md](方法说明_零基础版.md).
## Environment
The project is managed with `uv`. The Linux environment selects PyTorch's CUDA
13.0 wheel; the checked machine is an RTX 5070 Ti under Fedora WSL.
```bash
cd Q1
uv sync
uv run python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
```
`uv.lock` records the exact environment. The root data directory remains
outside the Python environment and is not copied into the package.
## Q1-E1: extract the 100 raw-video samples
```bash
cd Q1
uv run python -m q1.extract_features --output-dir outputs/q1_features
```
The first run downloads the pinned-by-name Hugging Face model revisions to the
local cache and the MediaPipe face model to `Q1/models/`. Use `--resume` to
continue an interrupted run. The supplied transcripts remain unchanged.
Each compressed NPZ keeps variable-length features without disk padding:
- Text: 768-d BERT vectors, one mean-pooled vector per supplied whitespace word.
- Audio: 25-d eGeMAPSv02 low-level descriptors from mono 16 kHz audio, with
frame-center timestamps.
- Vision: 192-d DeiT-Tiny CLS embeddings for every frame sampled at 5 fps;
optional 52-d MediaPipe face blendshapes have a separate validity mask.
- Alignment: CTC Viterbi word intervals in seconds from clip start, plus
word-level mean Audio and Vision features. Missing CTC word spans are
interpolated only as a flagged fallback (`word_alignment_valid=false`).
Outputs are written to `outputs/q1_features/`: one NPZ per original
`video_id/clip_id`, JSONL sample logs, a 100-row sample table, a 300-row
sample-by-modality table, a run manifest, and a typical-sample figure under
`reports/`. Missing faces do not remove the general Vision embedding; missing
modality values remain masked and are logged.
The extracted feature files and reports are approximately 10 MB before any
optional Q1 method-comparison outputs. Keep the final combined Q1/Q2/Q3
attachments under the problem's 50 MB cap.
## Q1-E0: sample coverage audit
```bash
cd Q1
uv run python -m q1.audit
```
The audit checks that the label workbook contains exactly 100 unique
`(video_id, clip_id)` pairs, resolves each pair to one MP4, counts the 37
`video_id` groups, probes each MP4's duration and audio/video streams, and checks
polarity labels against the sign of the continuous label. It writes
`outputs/audit/manifest.csv` and `outputs/audit/coverage_summary.json`.
## Required Q1 deliverables
1. Three-modality features generated from the original 100 videos and their
supplied transcripts. Keep every original row and label; if a stream needs a
mask or truncation, retain the source and record the rule.
2. Per-sample and per-modality records for effective duration, feature
dimension, valid length, padding rule, alignment granularity, and mapping
back to source time.
3. A 100-sample results table and a typical-sample view connecting transcript
text, aligned speech intervals, video frame times, and feature positions.
4. A run manifest with tool/model versions, parameters, logs, and commands.
5. Compressed feature files that fit the competition's 50 MB total-attachment
limit together with the Q2/Q3 code and result files.
The Attachment 1 media and labels are fixed input. The labels are for the
required inventory and later prediction probes; they are not a required training
target for the primary Q1 alignment model.
## Common feature and alignment contract
Each modality is supplied as a padded `SequenceBatch`:
- `features`: `[B, L_m, D_m]` feature values;
- `times`: `[B, L_m]` timestamps in seconds from clip start;
- `valid`: `[B, L_m]` boolean mask; padded positions are false.
The duration vector is `[B]`. Every method returns `AlignmentOutput` with
`weights[m]` shaped `[B, K, L_m]`. Rows sum to one and columns for padding stay
zero. `aligned[m]` is the weighted or learned representation on the shared
`K=50` grid. M1 takes one ordered `[word_count, 2]` forced-word-interval array
per sample; its 50 grid intervals partition transcript word order and preserve
the measured word times. M2 uses equal-duration windows.
M3 uses text-order slots as queries for Audio and Vision. M4 learns K shared
latent queries and attends separately to all three sequences. Both use the same
module dimensions and can be trained with `alignment_training_loss`; M1/M2 have
no training loss.
## Initial metrics
`q1.metrics` provides normalized expected-time trajectories, monotonicity
violation rate, normalized attention entropy, and shortest contiguous 80%
attention width. Entropy and width are diagnostics, not one-direction ranking
scores. A flat trajectory has zero backward violations too, so inspect its
time span as well; otherwise a collapsed alignment can look monotone. Grid-slot
retrieval is also available as a representation-consistency probe; because its
positive is defined by the shared grid index, it is not independent evidence
of temporal correctness. Use a human-labeled event subset for direct temporal
claims.
## Optional method-comparison rules
- Extract all three modalities once with the same extractor versions and keep
source timestamps, model IDs, and preprocessing parameters in the manifest.
- Split clips by `video_id`; never allow clips from one source video into both
training and validation folds.
- Fit learned M3/M4 models only on training folds. Freeze them before probes.
- Keep reconstruction and emotion probes identical across alignment methods.
- Do not use emotion labels in the primary alignment loss.
- Record all seeds, fold IDs, attention matrices, masks, and metrics alongside
each output.
## M1–M4 grouped comparison runner
Run the optional method comparison after feature extraction and the coverage
audit are complete:
```bash
cd Q1
uv run python -m q1.compare_methods
```
The runner selects CUDA when available, fits feature normalization on each
training fold only, and uses five grouped folds by `video_id`. M3/M4 share the
same masked-reconstruction, cross-modal contrastive, and temporal-monotonicity
objective; emotion labels are withheld from alignment training. It writes
fold checkpoints, out-of-fold alignment matrices, trajectories, retrieval and
reconstruction probes, and a frozen emotion probe under
`outputs/method_comparison/`. Retrieval uses same-grid positives, so it is a
representation-consistency check rather than independent temporal ground
truth. Direct human IoU/MATE scores require a separately annotated event set.
On Fedora WSL, PyTorch's Triton CUDA backward path also needs a system C
compiler and Python headers (`sudo dnf install gcc python3-devel`); Python
packages remain managed by `uv`.
## M3/M4 attention-collapse follow-up
After the first comparison showed nearly uniform M3/M4 attention, the staged
follow-up changes only learned alignment losses and probes. It reuses the saved
feature files and original video-grouped folds, and leaves M1/M2 untouched:
```bash
cd Q1
uv run python -m q1.train_alignment_variants
```
The runner executes v2-a (Span), v2-b (Span + far-slot Diversity), and v2-c
(Span + Diversity + weak Temporal Band) with the same folds and three seeds.
It records `C_row` during validation, tests within-clip temporal retrieval,
and compares frozen masked reconstruction with shuffled cross-modal slots.
Full checkpoints and per-sample alignment matrices are written under
`outputs/alignment_v2/`; use `outputs/alignment_v2/report_bundle/` for a
lightweight shareable summary. RoPE/Gaussian-bias experiments should wait until
the attention maps show time-progressing local bands.
## M3/M4 learnability diagnostics
Before running more architecture or positional-encoding sweeps, diagnose the
attention-collapse behavior with the small D0–D5 suite plus the D1-S synthetic
control:
```bash
cd Q1
uv run python -m q1.alignment_debug
uv run python -m q1.synthetic_alignment_sanity
uv run python -m q1.alignment_key_position_debug
uv run python -m q1.alignment_heldout_debug
uv run python -m q1.summarize_alignment_debug
```
D0/D1 overfit one representative clip. D2/D3 use all 100 clips for an
in-sample optimization diagnostic, with one seed. D1 compares M4 with and
without fixed sinusoidal positions; D2 uses a timestamp-derived Gaussian
attention target; D3 adds the same reconstruction and contrastive components
used in the prior learned methods. The Gaussian target is only a weak temporal
prior, not human alignment ground truth. The run records component losses and
gradients for attention Q/K and M4 latent slots, then saves per-sample metrics,
heatmaps, trajectories, and checkpoints under `outputs/alignment_debug/`.
The optional synthetic control uses only `[t, t², 1]` source features to verify
that M3/M4 can fit a known time band when position information is explicit.
The source-time diagnostic adds Fourier time identity to attention keys while
keeping the projected content as values; the held-out runner uses grouped fold
1 by default and fits feature normalization on its training videos only. Its
`--fold` and `--output-dir` options allow the same diagnostic to run on the
remaining grouped folds.
Summarize the existing run without retraining with the second command; it also
builds a compact `outputs/alignment_debug/report_bundle/` without checkpoints.
The diagnostics found nonzero gradients. D4/D5 show that M4 Audio/Vision can
form local bands with explicit time identity, while M3 and the M4 Text branch
still need work. Keep follow-up focused on those branches instead of broad
positional-encoding sweeps; see [RESULTS.md](RESULTS.md) for metrics and
interpretation.
## Same-slot vs shifted-time correspondence evaluation
Temporal attention bands only show where a slot looks. To check whether the
frozen content in the slot can match the other modalities, first create D5
checkpoints for all five `video_id`-grouped folds with
`q1.alignment_heldout_debug`, then run:
```bash
cd Q1
uv run python -m q1.correspondence_eval --device cuda --epochs 40 --seed 42
```
The evaluator applies the same train-only linear probe to M1, M2, and both
time-code settings of M3/M4. It reports within-video shifted-similarity curves,
same-video temporal retrieval, and matched-vs-shifted AUC on held-out videos.
The probe learns same-slot matching from training clips only; its held-out
scores indicate whether that matchability transfers, not whether semantic
alignment has independent ground truth. Outputs and grouped confidence
intervals are written to `outputs/correspondence_eval/`; see
[RESULTS.md](RESULTS.md) for the current results.
## M4 Shared Latent structural and functional re-evaluation
Once the five grouped D5 `M4_sourceTime` checkpoints exist, re-evaluate M4
without retraining its alignment model:
```bash
cd Q1
uv run python -m q1.m4_shared_latent_eval --device cuda --seed 42 --probe-epochs 40 --decoder-epochs 40 --shuffle-repeats 20
```
The evaluator checks self-Gram structure, derives six modality-to-modality
maps through the shared latent slots, measures physical-time trajectories,
cycle and triangle consistency, and runs held-out content-only correspondence,
within-clip retrieval, independent content-shuffle, and aligned-versus-shifted
reconstruction probes. It fits only these small evaluation probes on training
video groups; no emotion labels enter them. The reconstruction target is the
M4 attention-pooled content representation, not raw audio or video. Because
D5 used a timestamp-derived Gaussian prior, its attention weights may carry
time information indirectly even though the probe receives no explicit time
code or slot ID. Time structure and temporal maps therefore are not
independent semantic ground truth. Human event IoU/center error still requires
annotated events. CSVs, checkpoints, the run manifest, and figures are written
to `outputs/m4_shared_latent_eval/`; see [RESULTS.md](RESULTS.md) for the
five-fold results and interpretation.
## TSFA coarse-to-fine correspondence experiment
After the five grouped D5 checkpoints exist, freeze M4's source-time branch and
train a content-only local attention branch. The main variant searches within
`±0.10` normalized video time around each frozen M4 temporal center. A
multiply variant applies the M4 attention weights as an additional prior;
random-window and global-window controls test whether the temporal candidate
range matters. The original BERT text, eGeMAPS audio, and DeiT image features
stay fixed. Emotion labels are not used.
```bash
cd Q1
uv run python -m q1.tsfa_experiment --device cuda --seed 42
uv run python -m q1.tsfa_alignment_only_eval --device cuda --seed 42
uv run python -m q1.tsfa_emotion_probe --device cuda --seed 42
```
The first command fits one semantic branch and equal-seed correspondence probes
per fold, then evaluates all seven methods on held-out `video_id` groups. The
second pools the *same normalized raw source features* with each method's
alignment matrix before fitting equal probes. It separates source-position
selection from native model value/output projections. Both commands cover all
100 clips exactly once as held-out examples. Summary CSVs, per-clip metrics,
paired video-group confidence intervals, monotonicity/span/entropy diagnostics,
plots, manifests, and checkpoints are written to `outputs/tsfa/`; the small
`report_bundle/` contains shareable
summaries and figures. See [RESULTS.md](RESULTS.md) for the measured tradeoff
and interpretation limits. The third command evaluates the same frozen
representations with the grouped, five-segment logistic-regression and Ridge
probes; emotion labels are used only by these downstream probes.
With the math OOF predictions and fold assignment CSV available at their
workspace paths, compare TSFA against B0–B4 on the identical held-out samples
and bootstrap by source video:
```bash
uv run python -m q1.compare_emotion_probes --bootstrap-repeats 2000 --seed 42
```
The local Audio–Vision edge and modality-private residual experiment runs a
fixed 2×2 ablation with the same five grouped folds, temporal windows,
self-supervised objective, and emotion probes:
```bash
uv run python -m q1.tsfa_av_private_ablation --device cuda --seed 42
uv run python -m q1.compare_emotion_probes --tsfa-predictions outputs/tsfa_av_private_ablation/predictions.csv --candidate-method TSFA-T+Private --output-dir outputs/tsfa_av_private_ablation --output-stem private_vs_math --bootstrap-repeats 2000 --seed 42
```
The ablation files and figure are in `outputs/tsfa_av_private_ablation/`.
## TSFA-SPR shared/private factorization
Run the first grouped five-fold experiment with M4 temporal checkpoints frozen:
```bash
cd Q1
uv run python -m q1.tsfa_shared_private --device cuda --seed 42
```
The experiment pools each fold-standardized original text, audio, and vision
source stream with the matching frozen `M4_sourceTime` attention weights. A
shared MLP and three independent private MLPs learn from same-slot cross-modal
InfoNCE, within-modality orthogonality, and per-modality reconstruction. No
emotion labels enter representation training. The run also fits a train-fold
PCA control on raw private slots to match the 256-dimensional-per-slot SPR
control, then evaluates all representations with the existing grouped
LogisticRegression/Ridge probes. Outputs include the fold checkpoints,
training history, OOF predictions, grouped bootstrap contrasts, source and
reconstruction diagnostics, and figures under
`outputs/tsfa_shared_private/`.
If you only need to regenerate the derived source contrasts, assessment, and
plots from an existing run without retraining, use:
```bash
uv run python -m q1.tsfa_shared_private --finalize-existing
```
For the first seed's interpretation, diagnostic caveats, and next-step
decision, see [RESULTS.md](RESULTS.md#tsfa-spr时序共享私有分解) and
[`outputs/tsfa_shared_private/README.md`](outputs/tsfa_shared_private/README.md).
## Shared information definition comparison
Compare three concrete definitions of shared information using the same
fold-standardized original source features pooled by the frozen M4 temporal
checkpoints: Similarity (reuse the existing SPR weights), Correlation
(regularized MAXVAR/GCCA), and Predictability (cross-modal Ridge prediction
and its private residual). The Similarity network is loaded for inference and
is not retrained. Emotion labels are used only by the identical outer-fold
emotion probes.
```bash
cd Q1
uv run python -m q1.shared_definition_comparison --device cuda --seed 42
```
The run reuses the fixed five-fold `video_id` splits, compares same-slot,
shifted, and within-video-shuffled predictive controls, and reports shared
variance, residual predictability, dimension-matched 1,280D emotion probes,
private-source ablations, and paired video-group bootstrap contrasts against
RawPrivate-PCA, Similarity-SPR, TSFA+RawPrivate, and math B0/B4. The math
artifacts are read-only inputs. Tables, figures, fold assignments, and the run
manifest are written to `outputs/shared_definition_comparison/`; measured
results and limitations are in [RESULTS.md](RESULTS.md#shared-information-definition-similarity-vs-correlation-vs-predictability).
## Shared time coordinate with modality-private content
This follow-up uses the corrected definition: M4 supplies a frozen 50-slot
temporal coordinate, while BERT text, eGeMAPS audio, and DeiT vision retain
modality-specific content. Optional same-slot product/difference features are
evaluated only as task-specific interactions; they do not assert shared
semantics. Feature standardization, PCA, and emotion probes are fitted inside
each training fold. The math outputs are read-only historical controls.
```bash
cd Q1
uv run python -m q1.shared_time_private_content --device cuda --seed 42 --bootstrap-draws 2000
```
The run writes 19 new representation OOF probes, prior-result controls,
video-group bootstrap contrasts, temporal monitoring, fold/PCA audits, and nine
figures to `outputs/shared_time_private_content/`. See the dedicated
[results and seven-question interpretation](RESULTS.md#shared-time-private-content修正版验证).