建立分批同步基线(基础文件)
This commit is contained in:
@@ -0,0 +1,286 @@
|
||||
# Q1: feature extraction and temporal alignment
|
||||
|
||||
Q1's required output is a reproducible set of three-modality features extracted
|
||||
from all 100 raw videos in Attachment 1, with traceable timestamps and at least
|
||||
one inspectable text/audio/video example. The four-method comparison and
|
||||
learned probes are optional research extensions; they follow the extraction and
|
||||
submission artifacts rather than replacing them.
|
||||
|
||||
For a nontechnical introduction to the extraction steps and their limits, see [方法说明_零基础版.md](方法说明_零基础版.md).
|
||||
|
||||
## Environment
|
||||
|
||||
The project is managed with `uv`. The Linux environment selects PyTorch's CUDA
|
||||
13.0 wheel; the checked machine is an RTX 5070 Ti under Fedora WSL.
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv sync
|
||||
uv run python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
|
||||
```
|
||||
|
||||
`uv.lock` records the exact environment. The root data directory remains
|
||||
outside the Python environment and is not copied into the package.
|
||||
|
||||
## Q1-E1: extract the 100 raw-video samples
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.extract_features --output-dir outputs/q1_features
|
||||
```
|
||||
|
||||
The first run downloads the pinned-by-name Hugging Face model revisions to the
|
||||
local cache and the MediaPipe face model to `Q1/models/`. Use `--resume` to
|
||||
continue an interrupted run. The supplied transcripts remain unchanged.
|
||||
|
||||
Each compressed NPZ keeps variable-length features without disk padding:
|
||||
|
||||
- Text: 768-d BERT vectors, one mean-pooled vector per supplied whitespace word.
|
||||
- Audio: 25-d eGeMAPSv02 low-level descriptors from mono 16 kHz audio, with
|
||||
frame-center timestamps.
|
||||
- Vision: 192-d DeiT-Tiny CLS embeddings for every frame sampled at 5 fps;
|
||||
optional 52-d MediaPipe face blendshapes have a separate validity mask.
|
||||
- Alignment: CTC Viterbi word intervals in seconds from clip start, plus
|
||||
word-level mean Audio and Vision features. Missing CTC word spans are
|
||||
interpolated only as a flagged fallback (`word_alignment_valid=false`).
|
||||
|
||||
Outputs are written to `outputs/q1_features/`: one NPZ per original
|
||||
`video_id/clip_id`, JSONL sample logs, a 100-row sample table, a 300-row
|
||||
sample-by-modality table, a run manifest, and a typical-sample figure under
|
||||
`reports/`. Missing faces do not remove the general Vision embedding; missing
|
||||
modality values remain masked and are logged.
|
||||
|
||||
The extracted feature files and reports are approximately 10 MB before any
|
||||
optional Q1 method-comparison outputs. Keep the final combined Q1/Q2/Q3
|
||||
attachments under the problem's 50 MB cap.
|
||||
|
||||
## Q1-E0: sample coverage audit
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.audit
|
||||
```
|
||||
|
||||
The audit checks that the label workbook contains exactly 100 unique
|
||||
`(video_id, clip_id)` pairs, resolves each pair to one MP4, counts the 37
|
||||
`video_id` groups, probes each MP4's duration and audio/video streams, and checks
|
||||
polarity labels against the sign of the continuous label. It writes
|
||||
`outputs/audit/manifest.csv` and `outputs/audit/coverage_summary.json`.
|
||||
|
||||
## Required Q1 deliverables
|
||||
|
||||
1. Three-modality features generated from the original 100 videos and their
|
||||
supplied transcripts. Keep every original row and label; if a stream needs a
|
||||
mask or truncation, retain the source and record the rule.
|
||||
2. Per-sample and per-modality records for effective duration, feature
|
||||
dimension, valid length, padding rule, alignment granularity, and mapping
|
||||
back to source time.
|
||||
3. A 100-sample results table and a typical-sample view connecting transcript
|
||||
text, aligned speech intervals, video frame times, and feature positions.
|
||||
4. A run manifest with tool/model versions, parameters, logs, and commands.
|
||||
5. Compressed feature files that fit the competition's 50 MB total-attachment
|
||||
limit together with the Q2/Q3 code and result files.
|
||||
|
||||
The Attachment 1 media and labels are fixed input. The labels are for the
|
||||
required inventory and later prediction probes; they are not a required training
|
||||
target for the primary Q1 alignment model.
|
||||
|
||||
## Common feature and alignment contract
|
||||
|
||||
Each modality is supplied as a padded `SequenceBatch`:
|
||||
|
||||
- `features`: `[B, L_m, D_m]` feature values;
|
||||
- `times`: `[B, L_m]` timestamps in seconds from clip start;
|
||||
- `valid`: `[B, L_m]` boolean mask; padded positions are false.
|
||||
|
||||
The duration vector is `[B]`. Every method returns `AlignmentOutput` with
|
||||
`weights[m]` shaped `[B, K, L_m]`. Rows sum to one and columns for padding stay
|
||||
zero. `aligned[m]` is the weighted or learned representation on the shared
|
||||
`K=50` grid. M1 takes one ordered `[word_count, 2]` forced-word-interval array
|
||||
per sample; its 50 grid intervals partition transcript word order and preserve
|
||||
the measured word times. M2 uses equal-duration windows.
|
||||
|
||||
M3 uses text-order slots as queries for Audio and Vision. M4 learns K shared
|
||||
latent queries and attends separately to all three sequences. Both use the same
|
||||
module dimensions and can be trained with `alignment_training_loss`; M1/M2 have
|
||||
no training loss.
|
||||
|
||||
## Initial metrics
|
||||
|
||||
`q1.metrics` provides normalized expected-time trajectories, monotonicity
|
||||
violation rate, normalized attention entropy, and shortest contiguous 80%
|
||||
attention width. Entropy and width are diagnostics, not one-direction ranking
|
||||
scores. A flat trajectory has zero backward violations too, so inspect its
|
||||
time span as well; otherwise a collapsed alignment can look monotone. Grid-slot
|
||||
retrieval is also available as a representation-consistency probe; because its
|
||||
positive is defined by the shared grid index, it is not independent evidence
|
||||
of temporal correctness. Use a human-labeled event subset for direct temporal
|
||||
claims.
|
||||
|
||||
## Optional method-comparison rules
|
||||
|
||||
- Extract all three modalities once with the same extractor versions and keep
|
||||
source timestamps, model IDs, and preprocessing parameters in the manifest.
|
||||
- Split clips by `video_id`; never allow clips from one source video into both
|
||||
training and validation folds.
|
||||
- Fit learned M3/M4 models only on training folds. Freeze them before probes.
|
||||
- Keep reconstruction and emotion probes identical across alignment methods.
|
||||
- Do not use emotion labels in the primary alignment loss.
|
||||
- Record all seeds, fold IDs, attention matrices, masks, and metrics alongside
|
||||
each output.
|
||||
|
||||
## M1–M4 grouped comparison runner
|
||||
|
||||
Run the optional method comparison after feature extraction and the coverage
|
||||
audit are complete:
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.compare_methods
|
||||
```
|
||||
|
||||
The runner selects CUDA when available, fits feature normalization on each
|
||||
training fold only, and uses five grouped folds by `video_id`. M3/M4 share the
|
||||
same masked-reconstruction, cross-modal contrastive, and temporal-monotonicity
|
||||
objective; emotion labels are withheld from alignment training. It writes
|
||||
fold checkpoints, out-of-fold alignment matrices, trajectories, retrieval and
|
||||
reconstruction probes, and a frozen emotion probe under
|
||||
`outputs/method_comparison/`. Retrieval uses same-grid positives, so it is a
|
||||
representation-consistency check rather than independent temporal ground
|
||||
truth. Direct human IoU/MATE scores require a separately annotated event set.
|
||||
On Fedora WSL, PyTorch's Triton CUDA backward path also needs a system C
|
||||
compiler and Python headers (`sudo dnf install gcc python3-devel`); Python
|
||||
packages remain managed by `uv`.
|
||||
|
||||
## M3/M4 attention-collapse follow-up
|
||||
|
||||
After the first comparison showed nearly uniform M3/M4 attention, the staged
|
||||
follow-up changes only learned alignment losses and probes. It reuses the saved
|
||||
feature files and original video-grouped folds, and leaves M1/M2 untouched:
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.train_alignment_variants
|
||||
```
|
||||
|
||||
The runner executes v2-a (Span), v2-b (Span + far-slot Diversity), and v2-c
|
||||
(Span + Diversity + weak Temporal Band) with the same folds and three seeds.
|
||||
It records `C_row` during validation, tests within-clip temporal retrieval,
|
||||
and compares frozen masked reconstruction with shuffled cross-modal slots.
|
||||
Full checkpoints and per-sample alignment matrices are written under
|
||||
`outputs/alignment_v2/`; use `outputs/alignment_v2/report_bundle/` for a
|
||||
lightweight shareable summary. RoPE/Gaussian-bias experiments should wait until
|
||||
the attention maps show time-progressing local bands.
|
||||
|
||||
## M3/M4 learnability diagnostics
|
||||
|
||||
Before running more architecture or positional-encoding sweeps, diagnose the
|
||||
attention-collapse behavior with the small D0–D5 suite plus the D1-S synthetic
|
||||
control:
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.alignment_debug
|
||||
uv run python -m q1.synthetic_alignment_sanity
|
||||
uv run python -m q1.alignment_key_position_debug
|
||||
uv run python -m q1.alignment_heldout_debug
|
||||
uv run python -m q1.summarize_alignment_debug
|
||||
```
|
||||
|
||||
D0/D1 overfit one representative clip. D2/D3 use all 100 clips for an
|
||||
in-sample optimization diagnostic, with one seed. D1 compares M4 with and
|
||||
without fixed sinusoidal positions; D2 uses a timestamp-derived Gaussian
|
||||
attention target; D3 adds the same reconstruction and contrastive components
|
||||
used in the prior learned methods. The Gaussian target is only a weak temporal
|
||||
prior, not human alignment ground truth. The run records component losses and
|
||||
gradients for attention Q/K and M4 latent slots, then saves per-sample metrics,
|
||||
heatmaps, trajectories, and checkpoints under `outputs/alignment_debug/`.
|
||||
The optional synthetic control uses only `[t, t², 1]` source features to verify
|
||||
that M3/M4 can fit a known time band when position information is explicit.
|
||||
The source-time diagnostic adds Fourier time identity to attention keys while
|
||||
keeping the projected content as values; the held-out runner uses grouped fold
|
||||
1 by default and fits feature normalization on its training videos only. Its
|
||||
`--fold` and `--output-dir` options allow the same diagnostic to run on the
|
||||
remaining grouped folds.
|
||||
Summarize the existing run without retraining with the second command; it also
|
||||
builds a compact `outputs/alignment_debug/report_bundle/` without checkpoints.
|
||||
The diagnostics found nonzero gradients. D4/D5 show that M4 Audio/Vision can
|
||||
form local bands with explicit time identity, while M3 and the M4 Text branch
|
||||
still need work. Keep follow-up focused on those branches instead of broad
|
||||
positional-encoding sweeps; see [RESULTS.md](RESULTS.md) for metrics and
|
||||
interpretation.
|
||||
|
||||
## Same-slot vs shifted-time correspondence evaluation
|
||||
|
||||
Temporal attention bands only show where a slot looks. To check whether the
|
||||
frozen content in the slot can match the other modalities, first create D5
|
||||
checkpoints for all five `video_id`-grouped folds with
|
||||
`q1.alignment_heldout_debug`, then run:
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.correspondence_eval --device cuda --epochs 40 --seed 42
|
||||
```
|
||||
|
||||
The evaluator applies the same train-only linear probe to M1, M2, and both
|
||||
time-code settings of M3/M4. It reports within-video shifted-similarity curves,
|
||||
same-video temporal retrieval, and matched-vs-shifted AUC on held-out videos.
|
||||
The probe learns same-slot matching from training clips only; its held-out
|
||||
scores indicate whether that matchability transfers, not whether semantic
|
||||
alignment has independent ground truth. Outputs and grouped confidence
|
||||
intervals are written to `outputs/correspondence_eval/`; see
|
||||
[RESULTS.md](RESULTS.md) for the current results.
|
||||
|
||||
## M4 Shared Latent structural and functional re-evaluation
|
||||
|
||||
Once the five grouped D5 `M4_sourceTime` checkpoints exist, re-evaluate M4
|
||||
without retraining its alignment model:
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.m4_shared_latent_eval --device cuda --seed 42 --probe-epochs 40 --decoder-epochs 40 --shuffle-repeats 20
|
||||
```
|
||||
|
||||
The evaluator checks self-Gram structure, derives six modality-to-modality
|
||||
maps through the shared latent slots, measures physical-time trajectories,
|
||||
cycle and triangle consistency, and runs held-out content-only correspondence,
|
||||
within-clip retrieval, independent content-shuffle, and aligned-versus-shifted
|
||||
reconstruction probes. It fits only these small evaluation probes on training
|
||||
video groups; no emotion labels enter them. The reconstruction target is the
|
||||
M4 attention-pooled content representation, not raw audio or video. Because
|
||||
D5 used a timestamp-derived Gaussian prior, its attention weights may carry
|
||||
time information indirectly even though the probe receives no explicit time
|
||||
code or slot ID. Time structure and temporal maps therefore are not
|
||||
independent semantic ground truth. Human event IoU/center error still requires
|
||||
annotated events. CSVs, checkpoints, the run manifest, and figures are written
|
||||
to `outputs/m4_shared_latent_eval/`; see [RESULTS.md](RESULTS.md) for the
|
||||
five-fold results and interpretation.
|
||||
|
||||
## TSFA coarse-to-fine correspondence experiment
|
||||
|
||||
After the five grouped D5 checkpoints exist, freeze M4's source-time branch and
|
||||
train a content-only local attention branch. The main variant searches within
|
||||
`±0.10` normalized video time around each frozen M4 temporal center. A
|
||||
multiply variant applies the M4 attention weights as an additional prior;
|
||||
random-window and global-window controls test whether the temporal candidate
|
||||
range matters. The original BERT text, eGeMAPS audio, and DeiT image features
|
||||
stay fixed. Emotion labels are not used.
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.tsfa_experiment --device cuda --seed 42
|
||||
uv run python -m q1.tsfa_alignment_only_eval --device cuda --seed 42
|
||||
```
|
||||
|
||||
The first command fits one semantic branch and equal-seed correspondence probes
|
||||
per fold, then evaluates all seven methods on held-out `video_id` groups. The
|
||||
second pools the *same normalized raw source features* with each method's
|
||||
alignment matrix before fitting equal probes. It separates source-position
|
||||
selection from native model value/output projections. Both commands cover all
|
||||
100 clips exactly once as held-out examples. Summary CSVs, per-clip metrics,
|
||||
paired video-group confidence intervals, monotonicity/span/entropy diagnostics,
|
||||
plots, manifests, and checkpoints are written to `outputs/tsfa/`; the small
|
||||
`report_bundle/` contains shareable
|
||||
summaries and figures. See [RESULTS.md](RESULTS.md) for the measured tradeoff
|
||||
and interpretation limits.
|
||||
|
||||
Reference in New Issue
Block a user