建立分批同步基线(基础文件)
This commit is contained in:
+19
@@ -0,0 +1,19 @@
|
|||||||
|
/E题数据/
|
||||||
|
/.venv/
|
||||||
|
/.uv/
|
||||||
|
/.python-version
|
||||||
|
/uv.lock
|
||||||
|
/math/.venv/
|
||||||
|
/math/.uv/
|
||||||
|
/math/.python-version
|
||||||
|
/math/uv.lock
|
||||||
|
/math/.cache/
|
||||||
|
/math/cache/
|
||||||
|
/math/**/__pycache__/
|
||||||
|
/math/**/*.py[cod]
|
||||||
|
/Q1/.venv/
|
||||||
|
/Q1/models/
|
||||||
|
/Q1/__pycache__/
|
||||||
|
/Q1/outputs/
|
||||||
|
# TEMPORARY: output batches are force-added from the saved path lists
|
||||||
|
/deep_learning/Q1/outputs/
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
3.14.7
|
||||||
@@ -0,0 +1,286 @@
|
|||||||
|
# Q1: feature extraction and temporal alignment
|
||||||
|
|
||||||
|
Q1's required output is a reproducible set of three-modality features extracted
|
||||||
|
from all 100 raw videos in Attachment 1, with traceable timestamps and at least
|
||||||
|
one inspectable text/audio/video example. The four-method comparison and
|
||||||
|
learned probes are optional research extensions; they follow the extraction and
|
||||||
|
submission artifacts rather than replacing them.
|
||||||
|
|
||||||
|
For a nontechnical introduction to the extraction steps and their limits, see [方法说明_零基础版.md](方法说明_零基础版.md).
|
||||||
|
|
||||||
|
## Environment
|
||||||
|
|
||||||
|
The project is managed with `uv`. The Linux environment selects PyTorch's CUDA
|
||||||
|
13.0 wheel; the checked machine is an RTX 5070 Ti under Fedora WSL.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv sync
|
||||||
|
uv run python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
|
||||||
|
```
|
||||||
|
|
||||||
|
`uv.lock` records the exact environment. The root data directory remains
|
||||||
|
outside the Python environment and is not copied into the package.
|
||||||
|
|
||||||
|
## Q1-E1: extract the 100 raw-video samples
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.extract_features --output-dir outputs/q1_features
|
||||||
|
```
|
||||||
|
|
||||||
|
The first run downloads the pinned-by-name Hugging Face model revisions to the
|
||||||
|
local cache and the MediaPipe face model to `Q1/models/`. Use `--resume` to
|
||||||
|
continue an interrupted run. The supplied transcripts remain unchanged.
|
||||||
|
|
||||||
|
Each compressed NPZ keeps variable-length features without disk padding:
|
||||||
|
|
||||||
|
- Text: 768-d BERT vectors, one mean-pooled vector per supplied whitespace word.
|
||||||
|
- Audio: 25-d eGeMAPSv02 low-level descriptors from mono 16 kHz audio, with
|
||||||
|
frame-center timestamps.
|
||||||
|
- Vision: 192-d DeiT-Tiny CLS embeddings for every frame sampled at 5 fps;
|
||||||
|
optional 52-d MediaPipe face blendshapes have a separate validity mask.
|
||||||
|
- Alignment: CTC Viterbi word intervals in seconds from clip start, plus
|
||||||
|
word-level mean Audio and Vision features. Missing CTC word spans are
|
||||||
|
interpolated only as a flagged fallback (`word_alignment_valid=false`).
|
||||||
|
|
||||||
|
Outputs are written to `outputs/q1_features/`: one NPZ per original
|
||||||
|
`video_id/clip_id`, JSONL sample logs, a 100-row sample table, a 300-row
|
||||||
|
sample-by-modality table, a run manifest, and a typical-sample figure under
|
||||||
|
`reports/`. Missing faces do not remove the general Vision embedding; missing
|
||||||
|
modality values remain masked and are logged.
|
||||||
|
|
||||||
|
The extracted feature files and reports are approximately 10 MB before any
|
||||||
|
optional Q1 method-comparison outputs. Keep the final combined Q1/Q2/Q3
|
||||||
|
attachments under the problem's 50 MB cap.
|
||||||
|
|
||||||
|
## Q1-E0: sample coverage audit
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.audit
|
||||||
|
```
|
||||||
|
|
||||||
|
The audit checks that the label workbook contains exactly 100 unique
|
||||||
|
`(video_id, clip_id)` pairs, resolves each pair to one MP4, counts the 37
|
||||||
|
`video_id` groups, probes each MP4's duration and audio/video streams, and checks
|
||||||
|
polarity labels against the sign of the continuous label. It writes
|
||||||
|
`outputs/audit/manifest.csv` and `outputs/audit/coverage_summary.json`.
|
||||||
|
|
||||||
|
## Required Q1 deliverables
|
||||||
|
|
||||||
|
1. Three-modality features generated from the original 100 videos and their
|
||||||
|
supplied transcripts. Keep every original row and label; if a stream needs a
|
||||||
|
mask or truncation, retain the source and record the rule.
|
||||||
|
2. Per-sample and per-modality records for effective duration, feature
|
||||||
|
dimension, valid length, padding rule, alignment granularity, and mapping
|
||||||
|
back to source time.
|
||||||
|
3. A 100-sample results table and a typical-sample view connecting transcript
|
||||||
|
text, aligned speech intervals, video frame times, and feature positions.
|
||||||
|
4. A run manifest with tool/model versions, parameters, logs, and commands.
|
||||||
|
5. Compressed feature files that fit the competition's 50 MB total-attachment
|
||||||
|
limit together with the Q2/Q3 code and result files.
|
||||||
|
|
||||||
|
The Attachment 1 media and labels are fixed input. The labels are for the
|
||||||
|
required inventory and later prediction probes; they are not a required training
|
||||||
|
target for the primary Q1 alignment model.
|
||||||
|
|
||||||
|
## Common feature and alignment contract
|
||||||
|
|
||||||
|
Each modality is supplied as a padded `SequenceBatch`:
|
||||||
|
|
||||||
|
- `features`: `[B, L_m, D_m]` feature values;
|
||||||
|
- `times`: `[B, L_m]` timestamps in seconds from clip start;
|
||||||
|
- `valid`: `[B, L_m]` boolean mask; padded positions are false.
|
||||||
|
|
||||||
|
The duration vector is `[B]`. Every method returns `AlignmentOutput` with
|
||||||
|
`weights[m]` shaped `[B, K, L_m]`. Rows sum to one and columns for padding stay
|
||||||
|
zero. `aligned[m]` is the weighted or learned representation on the shared
|
||||||
|
`K=50` grid. M1 takes one ordered `[word_count, 2]` forced-word-interval array
|
||||||
|
per sample; its 50 grid intervals partition transcript word order and preserve
|
||||||
|
the measured word times. M2 uses equal-duration windows.
|
||||||
|
|
||||||
|
M3 uses text-order slots as queries for Audio and Vision. M4 learns K shared
|
||||||
|
latent queries and attends separately to all three sequences. Both use the same
|
||||||
|
module dimensions and can be trained with `alignment_training_loss`; M1/M2 have
|
||||||
|
no training loss.
|
||||||
|
|
||||||
|
## Initial metrics
|
||||||
|
|
||||||
|
`q1.metrics` provides normalized expected-time trajectories, monotonicity
|
||||||
|
violation rate, normalized attention entropy, and shortest contiguous 80%
|
||||||
|
attention width. Entropy and width are diagnostics, not one-direction ranking
|
||||||
|
scores. A flat trajectory has zero backward violations too, so inspect its
|
||||||
|
time span as well; otherwise a collapsed alignment can look monotone. Grid-slot
|
||||||
|
retrieval is also available as a representation-consistency probe; because its
|
||||||
|
positive is defined by the shared grid index, it is not independent evidence
|
||||||
|
of temporal correctness. Use a human-labeled event subset for direct temporal
|
||||||
|
claims.
|
||||||
|
|
||||||
|
## Optional method-comparison rules
|
||||||
|
|
||||||
|
- Extract all three modalities once with the same extractor versions and keep
|
||||||
|
source timestamps, model IDs, and preprocessing parameters in the manifest.
|
||||||
|
- Split clips by `video_id`; never allow clips from one source video into both
|
||||||
|
training and validation folds.
|
||||||
|
- Fit learned M3/M4 models only on training folds. Freeze them before probes.
|
||||||
|
- Keep reconstruction and emotion probes identical across alignment methods.
|
||||||
|
- Do not use emotion labels in the primary alignment loss.
|
||||||
|
- Record all seeds, fold IDs, attention matrices, masks, and metrics alongside
|
||||||
|
each output.
|
||||||
|
|
||||||
|
## M1–M4 grouped comparison runner
|
||||||
|
|
||||||
|
Run the optional method comparison after feature extraction and the coverage
|
||||||
|
audit are complete:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.compare_methods
|
||||||
|
```
|
||||||
|
|
||||||
|
The runner selects CUDA when available, fits feature normalization on each
|
||||||
|
training fold only, and uses five grouped folds by `video_id`. M3/M4 share the
|
||||||
|
same masked-reconstruction, cross-modal contrastive, and temporal-monotonicity
|
||||||
|
objective; emotion labels are withheld from alignment training. It writes
|
||||||
|
fold checkpoints, out-of-fold alignment matrices, trajectories, retrieval and
|
||||||
|
reconstruction probes, and a frozen emotion probe under
|
||||||
|
`outputs/method_comparison/`. Retrieval uses same-grid positives, so it is a
|
||||||
|
representation-consistency check rather than independent temporal ground
|
||||||
|
truth. Direct human IoU/MATE scores require a separately annotated event set.
|
||||||
|
On Fedora WSL, PyTorch's Triton CUDA backward path also needs a system C
|
||||||
|
compiler and Python headers (`sudo dnf install gcc python3-devel`); Python
|
||||||
|
packages remain managed by `uv`.
|
||||||
|
|
||||||
|
## M3/M4 attention-collapse follow-up
|
||||||
|
|
||||||
|
After the first comparison showed nearly uniform M3/M4 attention, the staged
|
||||||
|
follow-up changes only learned alignment losses and probes. It reuses the saved
|
||||||
|
feature files and original video-grouped folds, and leaves M1/M2 untouched:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.train_alignment_variants
|
||||||
|
```
|
||||||
|
|
||||||
|
The runner executes v2-a (Span), v2-b (Span + far-slot Diversity), and v2-c
|
||||||
|
(Span + Diversity + weak Temporal Band) with the same folds and three seeds.
|
||||||
|
It records `C_row` during validation, tests within-clip temporal retrieval,
|
||||||
|
and compares frozen masked reconstruction with shuffled cross-modal slots.
|
||||||
|
Full checkpoints and per-sample alignment matrices are written under
|
||||||
|
`outputs/alignment_v2/`; use `outputs/alignment_v2/report_bundle/` for a
|
||||||
|
lightweight shareable summary. RoPE/Gaussian-bias experiments should wait until
|
||||||
|
the attention maps show time-progressing local bands.
|
||||||
|
|
||||||
|
## M3/M4 learnability diagnostics
|
||||||
|
|
||||||
|
Before running more architecture or positional-encoding sweeps, diagnose the
|
||||||
|
attention-collapse behavior with the small D0–D5 suite plus the D1-S synthetic
|
||||||
|
control:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.alignment_debug
|
||||||
|
uv run python -m q1.synthetic_alignment_sanity
|
||||||
|
uv run python -m q1.alignment_key_position_debug
|
||||||
|
uv run python -m q1.alignment_heldout_debug
|
||||||
|
uv run python -m q1.summarize_alignment_debug
|
||||||
|
```
|
||||||
|
|
||||||
|
D0/D1 overfit one representative clip. D2/D3 use all 100 clips for an
|
||||||
|
in-sample optimization diagnostic, with one seed. D1 compares M4 with and
|
||||||
|
without fixed sinusoidal positions; D2 uses a timestamp-derived Gaussian
|
||||||
|
attention target; D3 adds the same reconstruction and contrastive components
|
||||||
|
used in the prior learned methods. The Gaussian target is only a weak temporal
|
||||||
|
prior, not human alignment ground truth. The run records component losses and
|
||||||
|
gradients for attention Q/K and M4 latent slots, then saves per-sample metrics,
|
||||||
|
heatmaps, trajectories, and checkpoints under `outputs/alignment_debug/`.
|
||||||
|
The optional synthetic control uses only `[t, t², 1]` source features to verify
|
||||||
|
that M3/M4 can fit a known time band when position information is explicit.
|
||||||
|
The source-time diagnostic adds Fourier time identity to attention keys while
|
||||||
|
keeping the projected content as values; the held-out runner uses grouped fold
|
||||||
|
1 by default and fits feature normalization on its training videos only. Its
|
||||||
|
`--fold` and `--output-dir` options allow the same diagnostic to run on the
|
||||||
|
remaining grouped folds.
|
||||||
|
Summarize the existing run without retraining with the second command; it also
|
||||||
|
builds a compact `outputs/alignment_debug/report_bundle/` without checkpoints.
|
||||||
|
The diagnostics found nonzero gradients. D4/D5 show that M4 Audio/Vision can
|
||||||
|
form local bands with explicit time identity, while M3 and the M4 Text branch
|
||||||
|
still need work. Keep follow-up focused on those branches instead of broad
|
||||||
|
positional-encoding sweeps; see [RESULTS.md](RESULTS.md) for metrics and
|
||||||
|
interpretation.
|
||||||
|
|
||||||
|
## Same-slot vs shifted-time correspondence evaluation
|
||||||
|
|
||||||
|
Temporal attention bands only show where a slot looks. To check whether the
|
||||||
|
frozen content in the slot can match the other modalities, first create D5
|
||||||
|
checkpoints for all five `video_id`-grouped folds with
|
||||||
|
`q1.alignment_heldout_debug`, then run:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.correspondence_eval --device cuda --epochs 40 --seed 42
|
||||||
|
```
|
||||||
|
|
||||||
|
The evaluator applies the same train-only linear probe to M1, M2, and both
|
||||||
|
time-code settings of M3/M4. It reports within-video shifted-similarity curves,
|
||||||
|
same-video temporal retrieval, and matched-vs-shifted AUC on held-out videos.
|
||||||
|
The probe learns same-slot matching from training clips only; its held-out
|
||||||
|
scores indicate whether that matchability transfers, not whether semantic
|
||||||
|
alignment has independent ground truth. Outputs and grouped confidence
|
||||||
|
intervals are written to `outputs/correspondence_eval/`; see
|
||||||
|
[RESULTS.md](RESULTS.md) for the current results.
|
||||||
|
|
||||||
|
## M4 Shared Latent structural and functional re-evaluation
|
||||||
|
|
||||||
|
Once the five grouped D5 `M4_sourceTime` checkpoints exist, re-evaluate M4
|
||||||
|
without retraining its alignment model:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.m4_shared_latent_eval --device cuda --seed 42 --probe-epochs 40 --decoder-epochs 40 --shuffle-repeats 20
|
||||||
|
```
|
||||||
|
|
||||||
|
The evaluator checks self-Gram structure, derives six modality-to-modality
|
||||||
|
maps through the shared latent slots, measures physical-time trajectories,
|
||||||
|
cycle and triangle consistency, and runs held-out content-only correspondence,
|
||||||
|
within-clip retrieval, independent content-shuffle, and aligned-versus-shifted
|
||||||
|
reconstruction probes. It fits only these small evaluation probes on training
|
||||||
|
video groups; no emotion labels enter them. The reconstruction target is the
|
||||||
|
M4 attention-pooled content representation, not raw audio or video. Because
|
||||||
|
D5 used a timestamp-derived Gaussian prior, its attention weights may carry
|
||||||
|
time information indirectly even though the probe receives no explicit time
|
||||||
|
code or slot ID. Time structure and temporal maps therefore are not
|
||||||
|
independent semantic ground truth. Human event IoU/center error still requires
|
||||||
|
annotated events. CSVs, checkpoints, the run manifest, and figures are written
|
||||||
|
to `outputs/m4_shared_latent_eval/`; see [RESULTS.md](RESULTS.md) for the
|
||||||
|
five-fold results and interpretation.
|
||||||
|
|
||||||
|
## TSFA coarse-to-fine correspondence experiment
|
||||||
|
|
||||||
|
After the five grouped D5 checkpoints exist, freeze M4's source-time branch and
|
||||||
|
train a content-only local attention branch. The main variant searches within
|
||||||
|
`±0.10` normalized video time around each frozen M4 temporal center. A
|
||||||
|
multiply variant applies the M4 attention weights as an additional prior;
|
||||||
|
random-window and global-window controls test whether the temporal candidate
|
||||||
|
range matters. The original BERT text, eGeMAPS audio, and DeiT image features
|
||||||
|
stay fixed. Emotion labels are not used.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.tsfa_experiment --device cuda --seed 42
|
||||||
|
uv run python -m q1.tsfa_alignment_only_eval --device cuda --seed 42
|
||||||
|
```
|
||||||
|
|
||||||
|
The first command fits one semantic branch and equal-seed correspondence probes
|
||||||
|
per fold, then evaluates all seven methods on held-out `video_id` groups. The
|
||||||
|
second pools the *same normalized raw source features* with each method's
|
||||||
|
alignment matrix before fitting equal probes. It separates source-position
|
||||||
|
selection from native model value/output projections. Both commands cover all
|
||||||
|
100 clips exactly once as held-out examples. Summary CSVs, per-clip metrics,
|
||||||
|
paired video-group confidence intervals, monotonicity/span/entropy diagnostics,
|
||||||
|
plots, manifests, and checkpoints are written to `outputs/tsfa/`; the small
|
||||||
|
`report_bundle/` contains shareable
|
||||||
|
summaries and figures. See [RESULTS.md](RESULTS.md) for the measured tradeoff
|
||||||
|
and interpretation limits.
|
||||||
|
|
||||||
@@ -0,0 +1,356 @@
|
|||||||
|
# Q1 特征提取结果
|
||||||
|
|
||||||
|
## 题面要求与完成范围
|
||||||
|
|
||||||
|
按题面第 2、3、5、6 页,Q1 首先要从附件 1 的原始 100 条视频建立可追溯的文本、语音、视觉特征,保持样本覆盖完整,并给出 100 条汇总表、至少一个文本—语音—视频帧对应样例、工具参数与复现信息。原始视频和标签没有被修改。
|
||||||
|
|
||||||
|
全量特征提取与对应性核验已经完成;随后也按 `video_id` 分组跑完了 M1–M4 的五折方法比较、重构探针与冻结情感探针。结果显示当前 M3/M4 训练目标会产生注意力塌缩,详见文末比较结果;因此没有宣称某个方法已经获得可靠的跨模态对齐。
|
||||||
|
|
||||||
|
## 数据覆盖审计
|
||||||
|
|
||||||
|
| 核验项 | 结果 |
|
||||||
|
| --- | ---: |
|
||||||
|
| 标签表样本 | 100 |
|
||||||
|
| 唯一 `(video_id, clip_id)` | 100 |
|
||||||
|
| 视频文件 | 100 |
|
||||||
|
| `video_id` 分组 | 37 |
|
||||||
|
| 音视频均可读取 | 100 |
|
||||||
|
| 重复、缺失、额外文件 | 0 |
|
||||||
|
| 标签极性不匹配 | 0 |
|
||||||
|
|
||||||
|
题面给出的片段时长范围下限为 2.648 秒。按解码视频帧时间统计,有两条原始视频短于该下限:`-mJ2ud6oKI8/6` 为 2.2356 秒,`-s9qJ7ATP7w/6` 为 2.4667 秒。两条均保留并正常提取,没有为满足区间而裁剪或删除。详见 [coverage_summary.json](outputs/audit/coverage_summary.json)。
|
||||||
|
|
||||||
|
## 提取结果
|
||||||
|
|
||||||
|
| 模态/特征 | 每位置维数 | 全部样本总长度 | 有效覆盖 |
|
||||||
|
| --- | ---: | ---: | ---: |
|
||||||
|
| 文本:BERT 词向量 | 768 | 1,932 个词 | 100/100 样本 |
|
||||||
|
| 音频:eGeMAPSv02 LLD | 25 | 77,261 帧 | 100/100 样本 |
|
||||||
|
| 视觉:DeiT-Tiny 帧 CLS 向量,5 fps | 192 | 3,948 帧 | 100/100 样本 |
|
||||||
|
| 补充视觉:MediaPipe 人脸 blendshape | 52 | 随人脸帧变化 | 87/100 样本检测到人脸 |
|
||||||
|
|
||||||
|
词级强制对齐共直接得到 1,920/1,932 个词的时间区间。另有 12 个词分布在 9 条样本中无法直接对齐;它们以单调的零宽度时间点插值,并在 `word_alignment_valid` 中标为 `false`。这些词的音频和视觉词级特征用最近的有效源帧回退,回退标志单独保存在 NPZ 中。
|
||||||
|
|
||||||
|
本次运行 100 条均无提取错误;13 条样本没有检测到人脸,但它们仍有有效的 DeiT 通用视觉帧嵌入。每条样本一个 NPZ,数值特征存为 float16,时间戳为 float32,掩码为 bool。文件中同时保留原始变长时间序列和按词级时间区间平均后的 Audio/Vision 特征,不做磁盘填充。
|
||||||
|
|
||||||
|
## 典型样本
|
||||||
|
|
||||||
|
样本 `-tPCytz4rww/12` 展示了 transcript 词序、CTC 强制对齐区间、eGeMAPS 时序特征,以及相同片段内的抽样视频帧。例图里的文本片段为 “and going to similar in, similar with bedrooms and bathrooms as well”。
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
## 产物与复现
|
||||||
|
|
||||||
|
- [100 条样本汇总表](outputs/q1_features/reports/sample_summary.csv)
|
||||||
|
- [100×3 样本—模态明细表](outputs/q1_features/reports/feature_summary_100x3.csv)
|
||||||
|
- [逐样本运行日志](outputs/q1_features/logs/samples.jsonl)
|
||||||
|
- [模型、版本、参数、哈希与运行命令](outputs/q1_features/reports/run_manifest.json)
|
||||||
|
- [100 个压缩特征文件](outputs/q1_features/features/)
|
||||||
|
|
||||||
|
100 个 NPZ 合计 9,261,634 字节;核心提取报告、日志与特征文件合计 10,460,565 字节。连同下面的轻量对比报告包仍低于题面 50 MB 上限;不要把完整训练检查点和全部逐样本对齐矩阵也放进提交包。
|
||||||
|
|
||||||
|
复现步骤(在 `deep_learning` 目录执行):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv sync
|
||||||
|
uv run python -m q1.audit
|
||||||
|
uv run python -m q1.extract_features --output-dir outputs/q1_features
|
||||||
|
```
|
||||||
|
|
||||||
|
环境为 Python 3.14.7、PyTorch 2.14.0+cu130、torchvision 0.29.0,运行设备为 NVIDIA GeForce RTX 5070 Ti。具体依赖锁定在 [uv.lock](uv.lock),模型修订号、MediaPipe 资产 SHA-256、特征规则和运行参数记录在 `run_manifest.json`。
|
||||||
|
|
||||||
|
## M1–M4 方法比较
|
||||||
|
|
||||||
|
### 实验协议
|
||||||
|
|
||||||
|
- M1 使用 CTC 强制词时间;M2 把每条视频分成 50 个等长时间窗;M3 以文本顺序槽查询音频和视频;M4 使用 50 个共享潜在槽分别查询三个模态。
|
||||||
|
- M3/M4 采用相同的自监督目标:连续块跨模态重构、跨模态对比损失和时间单调性损失。情绪标签不进入对齐模型训练。
|
||||||
|
- 100 条样本按 37 个 `video_id` 分成 5 折;同一原始视频的片段不会跨训练/验证集。M3/M4 各用 3 个随机种子,所有特征归一化参数只由当前训练折估计。
|
||||||
|
- 在每折验证集上冻结对齐,再训练相同结构的检索投影、连续块重构和轻量情绪 Probe。情绪 Probe 使用五段时间池化后的正则化线性分类器/回归器,适合小样本诊断,不代表大型数据集上的最终模型。
|
||||||
|
|
||||||
|
### 主要发现
|
||||||
|
|
||||||
|
| 方法 | Audio 熵 ↓ | Vision 熵 ↓ | Audio 时间跨度比例 ↑ | Vision 时间跨度比例 ↑ | T→A R@1 ↑ | Audio 重构 MAE ↓ | 情绪宏 F1 ↑ | 情绪 Pearson ↑ |
|
||||||
|
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||||
|
| M1 强制时间 | 0.358 | 0.021 | 0.894 | 0.896 | 0.0012 | 0.691 | 0.437 | 0.531 |
|
||||||
|
| M2 固定窗口 | 0.392 | 0.016 | 0.968 | 0.974 | 0.0014 | 0.662 | 0.343 | 0.508 |
|
||||||
|
| M3 文本锚定注意力 | 0.992 | 0.983 | -0.002 | -0.005 | 0.0197 | 0.384 | 0.438 | 0.546 |
|
||||||
|
| M4 共享潜轴 | 1.000 | 1.000 | 0.000 | 0.000 | 0.0071 | 0.378 | 0.414 | 0.507 |
|
||||||
|
|
||||||
|
时间跨度比例是每条样本“最后槽的期望时间减第一槽的期望时间”,再除以片段时长。接近 1 表示大致走过整段视频;接近 0 表示所有槽都落在同一时间附近。熵接近 1 表示注意力近似平均分布。检索使用训练折拟合的相同线性投影,正样本是同片段、同共享槽,因此只是表示一致性探针,不是人工对齐真值。
|
||||||
|
|
||||||
|
当前结果清楚暴露了学习型方法的问题:M3 的音频/视觉注意力接近均匀,时间跨度接近零;M4 的三种模态注意力几乎完全均匀。M4 的 MVR 虽然为 0,但这是因为平坦轨迹没有“向后走”;所以 MVR 必须与时间跨度和注意力熵一起看。重构误差更低也不能单独证明对齐更好,因为全局平均表示本身容易被重构。冻结情绪 Probe 的折间波动也较大,不能据此断言 M3 的细小优势稳定存在。
|
||||||
|
|
||||||
|
因此第一轮结论是:M1/M2 提供了可解释、覆盖完整的时间基线;v1 的 M3/M4 没有学到可信的细粒度跨模态对齐。没有人工事件时间标注,第一轮不报告 IoU/MATE,也不声称有直接对齐准确率。
|
||||||
|
|
||||||
|
### v2 对齐塌缩复查
|
||||||
|
|
||||||
|
按顺序只改变 M3/M4 的损失项,原始特征、提取器、M1/M2、五折 `video_id` 分组、三个随机种子和训练预算均保持不变。三阶段共训练 90 个 M3/M4 模型,使用 CUDA RTX 5070 Ti,完整运行耗时约 18 分钟。
|
||||||
|
|
||||||
|
| 阶段 | 新增约束 | 固定设置 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| v2-a | Span | 最小轨迹跨度 0.7,权重 5.0 |
|
||||||
|
| v2-b | Span + Diversity | 另加相隔至少 6 个槽的注意力行余弦相似度,权重 0.5 |
|
||||||
|
| v2-c | Span + Diversity + Band | 另加弱时间带约束,容许偏差 0.10,权重 10.0 |
|
||||||
|
|
||||||
|
所有阶段继续使用相同的重构、对比和单调性损失。M3 的时间带目标取文本词时间;M4 使用 50 个均匀槽中心。另记录 `C_row`,即不同槽位注意力行的平均余弦相似度;它越接近 1,槽位关注的位置越相似。
|
||||||
|
|
||||||
|
| 阶段/方法 | Audio 跨度 | Vision 跨度 | C_row Audio | C_row Vision | T→A R@1(±1 槽) | T→A 精确 R@1 | T→A MASE(槽) | Audio 重构增益 |
|
||||||
|
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||||
|
| 固定 M1 | 0.894 | 0.896 | 0.000 | 0.028 | 0.067 | 0.025 | 16.28 | 0.012 |
|
||||||
|
| 固定 M2 | 0.968 | 0.974 | 0.000 | 0.016 | 0.077 | 0.028 | 16.00 | 0.016 |
|
||||||
|
| v2-a M3 | 0.025 | -0.014 | 0.847 | 0.979 | 0.175 | 0.074 | 14.14 | 0.036 |
|
||||||
|
| v2-b M3 | 0.030 | -0.016 | 0.777 | 0.961 | 0.191 | 0.083 | 13.59 | 0.060 |
|
||||||
|
| v2-c M3 | 0.030 | -0.017 | 0.784 | 0.964 | 0.185 | 0.080 | 13.55 | 0.057 |
|
||||||
|
| v2-a M4 | 0.001 | -0.002 | 0.994 | 0.999 | 0.060 | 0.024 | 17.94 | 0.002 |
|
||||||
|
| v2-b M4 | 0.002 | -0.001 | 0.983 | 1.000 | 0.061 | 0.025 | 17.90 | 0.005 |
|
||||||
|
| v2-c M4 | 0.005 | -0.001 | 0.965 | 0.999 | 0.069 | 0.029 | 15.09 | 0.010 |
|
||||||
|
|
||||||
|
跨度是验证样本上的平均“末槽期望时间减首槽期望时间”除以片段时长;负值表示两端略有倒退,仍属于未覆盖时间轴。检索用训练折拟合的投影,在每条留出片段自己的 50 个槽中找匹配位置;R@1 允许误差 ±1 槽,精确 R@1 要求同槽,MASE 是 top-1 的平均绝对槽位误差。随机猜测时 ±1 槽 R@1 约为 0.06。重构增益为 `MAE_shuffled - MAE_aligned`,正值表示保留跨模态槽位关系后误差下降。各均值汇总了折、种子和样本,样本量不应被当作独立视频数。
|
||||||
|
|
||||||
|
结果表明,Diversity 让 M3 的 Audio 注意力行有所分化,Audio 重构增益也从 v2-a 的 0.036 增至 v2-b 的 0.060;但 M3 Audio 时间跨度仍只有约 0.03,Vision 跨度接近零且 `C_row` 仍高。M4 的弱 Band 约束只带来有限改善,Audio/Vision 注意力仍接近塌缩,重构增益也很小。典型样本图同样显示 M3 Audio/Vision 近似水平轨迹,M4 三模态轨迹近乎平坦。因此这些阶段没有达到“沿时间移动的局部带状对齐”,也不足以证明 M3/M4 学到了可靠对齐;本轮不继续 RoPE、Gaussian Bias 或 latent 长度消融。下一步应先检查学习型注意力为何难以利用时间监督,并调整对齐目标或优化策略,再用相同探针复核。
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
### 文件与复现
|
||||||
|
|
||||||
|
- [对比总表](outputs/method_comparison/comparison_summary.csv)
|
||||||
|
- [逐样本对齐指标](outputs/method_comparison/alignment_metrics.csv) 与 [按方法汇总](outputs/method_comparison/alignment_summary.csv)
|
||||||
|
- [检索探针](outputs/method_comparison/retrieval_summary.csv)、[重构探针](outputs/method_comparison/reconstruction_summary.csv)、[情绪 Probe](outputs/method_comparison/emotion_summary.csv)
|
||||||
|
- [折分清单](outputs/method_comparison/splits.json)、[训练记录](outputs/method_comparison/training_summary.csv)、[运行清单](outputs/method_comparison/run_manifest.json)
|
||||||
|
- [典型样本对齐热图](outputs/method_comparison/typical_alignment_heatmaps.png) 与 [时间轨迹图](outputs/method_comparison/typical_alignment_trajectories.png)
|
||||||
|
- [v2 三阶段对比总表](outputs/alignment_v2/v2_c/comparison_summary.csv)、[v2-c 注意力指标](outputs/alignment_v2/v2_c/alignment_summary.csv)、[v2-c 检索摘要](outputs/alignment_v2/v2_c/retrieval_summary.csv) 与 [v2-c 重构摘要](outputs/alignment_v2/v2_c/reconstruction_summary.csv)
|
||||||
|
- [v2-c 典型样本热图](outputs/alignment_v2/v2_c/typical_alignment_heatmaps.png)、[v2-c 时间轨迹图](outputs/alignment_v2/v2_c/typical_alignment_trajectories.png)、[逐 epoch 训练记录](outputs/alignment_v2/v2_c/training_history.csv) 与 [运行日志](outputs/alignment_v2/experiment.log)
|
||||||
|
- [v2 轻量结果包](outputs/alignment_v2/report_bundle/README.md),约 4.6 MB;不含检查点和逐样本矩阵
|
||||||
|
|
||||||
|
复现命令:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv sync
|
||||||
|
uv run python -m q1.compare_methods
|
||||||
|
uv run python -m q1.train_alignment_variants
|
||||||
|
uv run python -m q1.alignment_debug
|
||||||
|
uv run python -m q1.synthetic_alignment_sanity
|
||||||
|
uv run python -m q1.alignment_key_position_debug
|
||||||
|
uv run python -m q1.alignment_heldout_debug
|
||||||
|
uv run python -m q1.summarize_alignment_debug
|
||||||
|
```
|
||||||
|
|
||||||
|
本机的 GPU Triton 路径还需要系统 GCC 和 Python 开发头文件;Fedora WSL 可用 `sudo dnf install gcc python3-devel` 安装。完整方法比较目录约 119 MB,其中 30 个折/种子检查点约 53.5 MB;它们是本地研究复现材料,不应和必须提交的特征文件一起全部打包到 50 MB 附件中。可提交/分享的轻量结果包位于 [report_bundle](outputs/method_comparison/report_bundle/),约 1.18 MB;全量训练检查点和 800 个逐样本矩阵仍保存在其上一级目录。
|
||||||
|
|
||||||
|
v2 塌缩复查的完整材料位于 `outputs/alignment_v2/`,约 348 MB;其中检查点和 1,800 个逐样本矩阵只用于本地研究复现。分享时使用 [v2 report_bundle](outputs/alignment_v2/report_bundle/README.md),约 4.6 MB。完整比较运行清单在 [alignment_v2/run_manifest.json](outputs/alignment_v2/run_manifest.json)。
|
||||||
|
|
||||||
|
### M3/M4 可学习性诊断(D0–D5 与 D1-S)
|
||||||
|
|
||||||
|
v2 后先做小规模诊断,再决定是否增加位置编码或结构消融。本轮复用现有的 100 条特征,没有修改提取器、M1 或 M2;在 Fedora WSL 中通过 `uv` 使用 RTX 5070 Ti 运行,耗时 103.7 秒,环境为 Python 3.14.7、PyTorch 2.14.0+cu130。
|
||||||
|
|
||||||
|
| 编号 | 数据与目标 | M4 位置编码 | 目的 |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| D0 | 单样本 `-tPCytz4rww/12`,`5 × span + 10 × band`,1,000 步 | 无 | 检查时间损失是否能推动注意力 |
|
||||||
|
| D1 | 同一样本,仅 Gaussian target KL,1,000 步 | 无 / 固定正弦位置编码 | 检查直接注意力监督与槽位对称性 |
|
||||||
|
| D1-S | 只用合成时间输入 `[t, t², 1]`,Gaussian target KL,1,000 步 | M4 无 PE / 有 PE | 极简验证注意力实现能否拟合已知时间带 |
|
||||||
|
| D2 | 全部 100 条,仅 Gaussian target KL,500 步 | 固定正弦位置编码 | 检查全数据下直接时间目标是否可优化 |
|
||||||
|
| D3 | 全部 100 条,Gaussian KL + 遮挡重构 + 跨模态对比,500 步 | 固定正弦位置编码 | 检查内容损失是否干扰时间监督 |
|
||||||
|
| D4 | 单样本 Gaussian KL,1,000 步;对照显式时间身份 | M4 保留 absolute PE | 检查真实特征加入时间码后能否形成局部带 |
|
||||||
|
| D5 | 按 `video_id` 分组 fold 1:80 条训练、20 条留出,纯 Gaussian KL,500 步 | 对照无时间码 / 有时间码 | 检查时间码能否迁移到未见原视频 |
|
||||||
|
|
||||||
|
Gaussian target 是根据时间戳构造的弱时间先验,归一化时间上的 `sigma=0.10`;它**不是人工对齐真值**。D0/D1 是单样本过拟合诊断;D2/D3 在相同 100 条数据上训练并评价、每种方法仅一个随机种子,因此只反映优化情况,不表示泛化表现。每 20 步记录损失分量以及注意力 Q/K 投影和 M4 latent slots 的梯度。
|
||||||
|
|
||||||
|
| 诊断 | 结果 |
|
||||||
|
| --- | --- |
|
||||||
|
| D0 时间损失 | M3/M4 的最终 span loss 均为 0;band loss 分别为 `1.28e-5`、`7.10e-5`。损失对 Q/K 有梯度,M4 slots 也有非零梯度。 |
|
||||||
|
| D1 真实特征的逐模态 KL | M3:Audio `0.5254`,Vision `0.0001`;M4 无 PE:Text/Audio/Vision `0.000009/0.5496/0.000075`;M4 有 PE:`0.000026/0.5452/0.000065`。总体均值掩盖了 Audio 目标拟合失败,PE 几乎没有改善 Audio。 |
|
||||||
|
| D1-S 合成时间输入 | M3 Audio/Vision KL `0.000098/0.000103`、轨迹跨度 `0.836/0.840`;M4 无 PE 的三模态 KL 均约 `0.000025–0.000032`,有 PE 约 `0.000024–0.000035`。轨迹跨度约 `0.835–0.845`,中心时间误差约 `0.012–0.014`。 |
|
||||||
|
| D2 全量训练最终 KL | M3 `0.5194`;M4 `0.6222`。Q/K 梯度非零,M4 slots 梯度也非零。 |
|
||||||
|
| D3 最终损失 | M3:total `2.3617`、align KL `0.8575`、reconstruction `0.0652`、contrastive `1.4390`。M4:total `2.5226`、align KL `0.7414`、reconstruction `0.0345`、contrastive `1.7467`。Q/K 与 M4 slots 梯度非零。 |
|
||||||
|
| D4 单样本时间码 | Audio KL:M3 `0.5254 → 0.0012`,M4 `0.5452 → 0.0011`;Audio 轨迹跨度:M3 `0.426 → 0.803`,M4 `0.439 → 0.824`。两者 Vision KL 均低于 `0.00013`。 |
|
||||||
|
| D5 留出 20 条、6 个未见 video_id | M4 时间码版 Audio/Vision KL `0.0035/0.0135`、跨度 `0.814/0.822`、中心误差 `0.017/0.017`;无时间码版 KL `0.994/1.359`、跨度 `0.010/-0.077`。M3 时间码版有改善,但留出 KL 仍为 `0.491/0.479`、跨度 `0.716/0.712`、中心误差约 `0.09`。 |
|
||||||
|
|
||||||
|
KL 的取平均模态数不同:M3 只对 Audio/Vision 计算 Gaussian target,M4 对 Text/Audio/Vision 都计算。因此 D1–D3 的 KL 数值用于看各自能否下降,不能直接作为 M3 对 M4 的胜负排名;D1 更应看逐模态值,而不是整体均值。
|
||||||
|
|
||||||
|
D1-S 使用实际 M3/M4 注意力层,但输入不含真实内容,只含合成时间坐标。它把 Gaussian 时间带拟合到约 `1e-4` KL,说明这套注意力计算和优化在输入携带时间位置时可以学出局部带状结构。结合 D1 真实特征上的 Audio 失败,问题更可能在真实 Audio 表示/时间线索与注意力目标的衔接,而非完全断开的 Q/K 梯度。代码检查也确认当前真实 M3/M4 注意力把投影后的源特征直接作为 key/value,没有显式把源时间戳编码进 key;这使“Audio key 缺少足够时间身份”成为优先验证的假设,但还不是已证实的唯一原因。
|
||||||
|
|
||||||
|
D4 为 source key 加固定 Fourier 时间码,value 保留投影后的真实特征;M3 query 同时加文本词时间中心码,M4 query 保留 absolute PE 并加对应网格时间码。真实单样本上两者都形成局部带。首次 D5 只跑了 fold 1:训练集 80 条、31 个 `video_id`,留出集 20 条、6 个 `video_id`,两侧无组重叠;归一化参数只用训练集拟合。该折支持 M4 的 Audio/Vision 时间带可迁移到未见视频;M3 优于无时间码基线,但仍有明显泛化差距。M4 Text 分支也仍弱,留出平均 Gaussian KL 为 `0.477`。后续新增的五折 D5 汇总见“Q1-Correspondence Evaluation”。D5 仍只有一个训练随机种子,Gaussian target 又来自源时间戳,因此这不是人工对齐准确率或多种子稳定性结论。
|
||||||
|
|
||||||
|
汇总 100 条样本的逐样本诊断,D2 的 M3 Audio 注意力仍较分散(归一化熵 `0.965`,平均轨迹跨度 `0.152`);Vision 更随时间变化(熵 `0.781`,跨度 `0.601`),但热图仍未清晰贴合目标。D2 的 M4 Audio 几乎均匀(熵 `0.997`,跨度 `0.011`),Vision 时间覆盖也有限(熵 `0.937`,跨度 `0.199`)。D3 加入重构和对比损失后,Audio 塌缩更明显:M3 熵/跨度为 `0.995/0.016`,M4 为 `0.999/0.002`;M4 Vision 跨度也从 D2 的 `0.199` 降至 `0.061`。因此重构误差降低不能单独视为对齐改善。
|
||||||
|
|
||||||
|
这组诊断排除了“注意力计算图完全断开或没有梯度”这一简单解释:D0 时间损失可优化,D1-S 可拟合合成时间带,全数据训练时 Q/K(以及 M4 slots)也有非零梯度。D4/D5 显示显式时间身份能解决 M4 Audio/Vision 的塌缩,但 M3 和 M4 Text 还未达到同样水平。后续优先检查 M3 文本查询时间身份和 M4 Text 分支,再考虑其他结构;暂不恢复大规模 RoPE 或 latent 长度扫参。可以先在其余 grouped folds 上复核 M4 音视频,再单独改进 M3/Text 时间条件化。
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
诊断产物见 [逐方法汇总指标](outputs/alignment_debug/metric_summary.csv)、[训练及梯度汇总](outputs/alignment_debug/debug_summary.csv)、[运行清单](outputs/alignment_debug/run_manifest.json)、[合成控制指标](outputs/alignment_debug/synthetic/metrics.csv)、[D4 时间码指标](outputs/alignment_debug/source_time/per_sample_metrics.csv)、[D5 留出汇总](outputs/alignment_debug/heldout/heldout_summary.csv) 和 [D5 运行清单](outputs/alignment_debug/heldout/run_manifest.json)。可分享的轻量结果包在 [report_bundle](outputs/alignment_debug/report_bundle/README.md);逐方法训练历史、指标、热图、轨迹及检查点保存在 `outputs/alignment_debug/`。
|
||||||
|
|
||||||
|
### Q1-Correspondence Evaluation:同槽与时间错位
|
||||||
|
|
||||||
|
为区分“注意力找到了时间”与“同一槽中的模态内容确实可匹配”,新增冻结特征上的对应性探针。M1、M2、M3、M4 共六个设置(M3/M4 各含有/无 source-time code)都使用同一套评估。每个方法的线性投影只在训练折上用同一片段内的同槽正样本和错槽负样本训练;再在完全留出的 `video_id` 上评价。投影维度为 64,训练 40 个 epoch,不使用情绪标签,也不按留出分数挑选模型。M3/M4 的 D5 时间定位模型在 5 个既有分组折上各训练 500 步、随机种子 42;每折 80 条训练片段和 20 条留出片段,100 条留出预测合计覆盖 37 个 `video_id`。
|
||||||
|
|
||||||
|
三项核心测量为:
|
||||||
|
|
||||||
|
1. **时间偏移曲线**:计算 `sim(z_i^m, z_(i+Δ)^n)`,`Δ=-10…10`;若同槽内容关系具有迁移性,曲线应在零偏移附近出现峰。
|
||||||
|
2. **片段内时序检索**:每个查询只能在同一留出片段的 50 个槽中检索,报告精确 R@1、±1 槽 R@1 和 MASE。
|
||||||
|
3. **匹配/错位 AUC**:正样本为同槽配对,负样本来自同一片段且 `|i-j|>2` 的错槽配对。AUC `0.5` 表示无法区分。表中置信区间以 `video_id` 为单位做 2,000 次 bootstrap,避免把同视频片段或时间槽当成独立样本。
|
||||||
|
|
||||||
|
| 冻结表示 | T–A AUC(95% CI) | T–V AUC | A–V AUC | T→A 精确 R@1 | T→A ±1 R@1 | T→A MASE(槽) |
|
||||||
|
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||||
|
| M1 强制时间 | 0.524 [0.510, 0.538] | 0.499 | 0.499 | 0.028 | 0.075 | 16.35 |
|
||||||
|
| M2 固定窗口 | 0.531 [0.518, 0.545] | 0.499 | 0.503 | 0.029 | 0.078 | 15.95 |
|
||||||
|
| M3 无时间码 | **0.728 [0.706, 0.751]** | 0.531 | 0.512 | **0.135** | **0.275** | **9.94** |
|
||||||
|
| M3 有时间码 | 0.542 [0.527, 0.557] | 0.500 | 0.505 | 0.036 | 0.088 | 16.60 |
|
||||||
|
| M4 无时间码 | 0.621 [0.602, 0.641] | 0.560 | 0.541 | 0.046 | 0.114 | 11.21 |
|
||||||
|
| M4 有时间码 | 0.546 [0.517, 0.573] | 0.490 | 0.500 | 0.029 | 0.076 | 17.11 |
|
||||||
|
|
||||||
|
50 槽下,随机检索参考值为精确 R@1 `0.020`、±1 R@1 `0.059`、MASE `16.66`。最清晰的偏移曲线峰出现在 **M3 无时间码的 T–A**:同槽相似度高于错开位置,T→A 检索也优于随机参考。M4 无时间码在三组配对上都有高于机会水平的 AUC,但效应较小;M1/M2 的 T–A 也有较弱的同槽信号。M3/M4 有时间码的结果没有显示同等强度的内容匹配,尤其 M4 有时间码的 T–V 与 A–V 接近机会水平。
|
||||||
|
|
||||||
|
这个结果必须和时间定位指标一起读。下表是五折 D5 验证片段按 `video_id` 宏平均的时间跨度;置信区间也按原视频分组。KL 越低越符合弱 Gaussian 时间目标。
|
||||||
|
|
||||||
|
| D5 变体 | Audio 跨度(95% CI) | Vision 跨度(95% CI) | Audio KL | Vision KL |
|
||||||
|
| --- | ---: | ---: | ---: | ---: |
|
||||||
|
| M3 无时间码 | 0.007 [-0.005, 0.018] | -0.052 [-0.092, -0.017] | 1.270 | 1.931 |
|
||||||
|
| M3 有时间码 | 0.708 [0.664, 0.738] | 0.710 [0.678, 0.737] | 0.618 | 0.592 |
|
||||||
|
| M4 无时间码 | 0.006 [0.001, 0.010] | -0.068 [-0.103, -0.033] | 1.013 | 1.550 |
|
||||||
|
| M4 有时间码 | 0.816 [0.815, 0.817] | 0.822 [0.819, 0.824] | 0.004 | 0.022 |
|
||||||
|
|
||||||
|
M4 有时间码的 Audio/Vision 时间带清晰,但其 Text 分支 Gaussian KL 仍为 `0.608`,T–A/T–V/A–V 内容对应性也弱。反过来,M3 无时间码的 Audio/Vision 时间轴没有覆盖整段片段,却在 T–A 匹配探针上表现最好。M4 无时间码的 Audio/Vision 轨迹同样塌缩,但三组 AUC 均有一定高于机会水平的统计信号。也就是说,**时间定位与跨模态可匹配性确实是两件不同的事**:零偏移匹配不等于整段视频已经正确对齐;时间带漂亮也不保证槽内内容可匹配。
|
||||||
|
|
||||||
|
因此目前最稳妥的结论是:M3 无时间码的文本—音频表示具有可迁移的同槽可匹配性,但它没有形成完整时间覆盖;M4 有时间码建立了更好的 Audio/Vision 时间定位,却还没有证明三模态内容对应;M4 无时间码的内容探针结果值得跟进,但其时间轨迹塌缩,不能称为完整时序对齐。投影头本身使用了训练视频的同槽标签,所以这些结果是“冻结表示可被简单探针匹配”的证据,不是无监督模型独立发现的真值,也不证明模态间语义相同。还需要多随机种子和人工事件标注确认。
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
逐片段结果、视频组 bootstrap 汇总、偏移曲线数据、五折 D5 时间定位汇总、探针训练记录、探针检查点和运行清单分别保存在 [correspondence_clip_metrics.csv](outputs/correspondence_eval/correspondence_clip_metrics.csv)、[metric_summary.csv](outputs/correspondence_eval/metric_summary.csv)、[shift_curve_summary.csv](outputs/correspondence_eval/shift_curve_summary.csv)、[time_localization_summary.csv](outputs/correspondence_eval/time_localization_summary.csv)、[probe_training_history.csv](outputs/correspondence_eval/probe_training_history.csv)、[probe_checkpoints.pt](outputs/correspondence_eval/probe_checkpoints.pt) 和 [run_manifest.json](outputs/correspondence_eval/run_manifest.json)。
|
||||||
|
|
||||||
|
完整复现(Fedora WSL,Python 环境由 `uv` 管理):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.alignment_heldout_debug --device cuda --seed 42 --fold 1 --output-dir outputs/alignment_debug/heldout
|
||||||
|
uv run python -m q1.alignment_heldout_debug --device cuda --seed 42 --fold 2 --output-dir outputs/alignment_debug/heldout/fold_02
|
||||||
|
uv run python -m q1.alignment_heldout_debug --device cuda --seed 42 --fold 3 --output-dir outputs/alignment_debug/heldout/fold_03
|
||||||
|
uv run python -m q1.alignment_heldout_debug --device cuda --seed 42 --fold 4 --output-dir outputs/alignment_debug/heldout/fold_04
|
||||||
|
uv run python -m q1.alignment_heldout_debug --device cuda --seed 42 --fold 5 --output-dir outputs/alignment_debug/heldout/fold_05
|
||||||
|
uv run python -m q1.correspondence_eval --device cuda --epochs 40 --seed 42
|
||||||
|
```
|
||||||
|
|
||||||
|
### M4 Shared Latent Timeline:结构与功能复核
|
||||||
|
|
||||||
|
M4 的判断不再以三张注意力图是否像对角线为依据,而分三层复核:先比较同一模态中 latent slots 的注意力相似度 (G^m=\hat A^m\hat A^{mT}),再由 (C^{m\to n}=\operatorname{ColNorm}(A^m)^T A^n) 生成模态到模态的转移图,最后用只看聚合内容的对应探针、内容打乱和错位重构检查这些关系是否有功能。此轮复用了五折 D5 `M4_sourceTime` 检查点,没有重新训练 M4 对齐模型。内容对应投影与小型重构器只在各折训练 video_id 上训练,再在留出组上打分;不使用情绪标签。全部 100 个片段各作为一次留出样本,分组覆盖 37 个 video_id。
|
||||||
|
|
||||||
|
结构指标按 video_id 做宏平均;95% 区间使用 2,000 次视频组 bootstrap。`near_similarity` 按定义包含自身对角线;`near_similarity_offdiag` 和 `d_self_offdiag` 排除自身,以免自相似度 1 掩盖邻近 slot 之间的结构。
|
||||||
|
|
||||||
|
| 模态 | 近邻相似(排除自身)↑ | 远邻相似 ↓ | (D_{self})(排除自身)↑ | Gram 目标误差 ↓ | 远槽泄漏 ↓ |
|
||||||
|
| --- | ---: | ---: | ---: | ---: | ---: |
|
||||||
|
| Text | 0.985 | 0.122 | 0.863 | 0.733 | 0.192 |
|
||||||
|
| Audio | 0.983 | 0.069 | 0.914 | 0.544 | 0.126 |
|
||||||
|
| Vision | 0.983 | 0.073 | 0.910 | 0.543 | 0.133 |
|
||||||
|
|
||||||
|
三种模态的 Gram 矩阵都有清晰的局部带:相邻 slots 会看相似区域,远距离 slots 的相似度明显更低。这说明 latent timeline 在各模态内部形成了平滑的时间结构;Text 的远槽相似和 Gram 目标误差仍高于 Audio/Vision。
|
||||||
|
|
||||||
|
| Pairwise 方向 | 物理时间 MAE ↓ | 时间相关 ↑ |
|
||||||
|
| --- | ---: | ---: |
|
||||||
|
| Text → Audio | 0.0841 | 0.928 |
|
||||||
|
| Text → Vision | 0.0843 | 0.928 |
|
||||||
|
| Audio → Vision | 0.0248 | 0.996 |
|
||||||
|
|
||||||
|
时间 MAE 使用归一化片段时间,描述由 pairwise map 得到的期望时间是否接近;它不是语义或事件级对齐准确率。完整六个方向见 [pairwise_summary.csv](outputs/m4_shared_latent_eval/pairwise_summary.csv)。回环到 band target 的相对误差分别为 Text–Audio–Text `0.865`、Text–Vision–Text `0.864`、Audio–Vision–Audio `0.719`;对应回环时间 MAE 为 `0.096`、`0.099`、`0.050`。Triangle (T\to A\to V) 与 (T\to V) 的相对 Frobenius 残差为 `0.226`。所以时间轨迹较一致,但 cycle 回环还没有恢复到期望的局部时间带,pairwise 结构尚未完全自洽。
|
||||||
|
|
||||||
|
功能验证的结果进一步收紧了结论。Content-only probe 输入仅为 M4 的注意力加权 Value 表示;不把 latent query、位置向量或 slot 编号拼入 probe。需要注意,source-time code 仍会影响注意力权重 (A^m),所以时间信息可能通过“选中了哪些 Value”间接保留;内容 shuffle 用来检验槽内内容顺序是否贡献了可匹配信号。探针使用训练折的同槽正例训练,留出时的负例是同片段中相隔超过 2 槽的位置。
|
||||||
|
|
||||||
|
| 配对 | 匹配/错位 AUC(95% CI) | 左到右精确 R@1 | ±1 槽 R@1 | MASE(槽) |
|
||||||
|
| --- | ---: | ---: | ---: | ---: |
|
||||||
|
| Text → Audio | **0.616 [0.583, 0.650]** | 0.037 | 0.108 | 13.64 |
|
||||||
|
| Text → Vision | 0.504 [0.494, 0.514] | 0.025 | 0.069 | 17.83 |
|
||||||
|
| Audio → Vision | 0.513 [0.500, 0.530] | 0.024 | 0.062 | 17.66 |
|
||||||
|
|
||||||
|
50 槽随机检索的参考值为精确 R@1 `0.020`、±1 R@1 `0.059`、MASE `16.66`。Text–Audio 的同槽相似曲线在零偏移处达到峰值,AUC 也高于机会水平;打乱槽内内容后 AUC 从 `0.616` 降至 `0.500`,说明这项有限信号依赖实际内容顺序。Text–Vision 与 Audio–Vision 的 AUC 均接近 `0.5`,打乱前后没有实质差别。因此内容层证据只支持较弱的 Text–Audio 对应,没有支持完整三模态内容对齐。
|
||||||
|
|
||||||
|
Aligned-vs-shifted reconstruction 的目标是另一个模态的 M4 聚合内容向量,而不是原始波形或视频像素;每次只将一个输入伙伴错开,且比较相同目标槽。下表为绝对位移 10 槽时 `MAE_shifted - MAE_aligned` 的估计和视频组 95% 区间:
|
||||||
|
|
||||||
|
| 重构目标 | 增益 (G(10))(95% CI) |
|
||||||
|
| --- | ---: |
|
||||||
|
| Audio | 0.000049 [-0.000150, 0.000260] |
|
||||||
|
| Vision | 0.000176 [-0.000229, 0.000624] |
|
||||||
|
| Text | 0.000675 [-0.000249, 0.001793] |
|
||||||
|
|
||||||
|
三个区间都包含 0;在 `|Δ|=1,2,5` 时同样没有排除 0。错位越远时 Text/Vision 目标的增益均值略有上升,但这项探针尚未给出稳定的正向重构证据。
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
**当前结论:** M4 有明显的 latent 自身时间结构,且 Audio–Vision pair map 在物理时间上接近一致;这与 D5 的 source-time 先验结果相符。可是 cycle band 误差仍大,内容对应主要只在 Text–Audio 上表现为中等信号,Text–Vision/Audio–Vision 接近机会水平,错位重构增益的区间包含 0。由此不能写成“M4 已学到可信的三模态共享语义对齐”。更稳妥的表述是:它建立了时间定位和部分 Text–Audio 内容可匹配性,但现有结果尚未证明完整、可靠的三模态跨模态对应。source-time D5 检查点只用一个对齐训练种子,而且训练中的 Gaussian 时间目标来自时间戳;本轮 bootstrap 只描述视频组采样不确定性,不覆盖训练种子不确定性。人工事件 IoU/中心误差仍需先标注 10–15 条视频的事件区间。
|
||||||
|
|
||||||
|
逐片段数据、汇总、探针训练记录和权重,以及可直接复查的典型样本矩阵保存在 [outputs/m4_shared_latent_eval](outputs/m4_shared_latent_eval/)。关键文件包括 [self_structure_summary.csv](outputs/m4_shared_latent_eval/self_structure_summary.csv)、[pairwise_summary.csv](outputs/m4_shared_latent_eval/pairwise_summary.csv)、[cycle_triangle_summary.csv](outputs/m4_shared_latent_eval/cycle_triangle_summary.csv)、[content_only_metrics_summary.csv](outputs/m4_shared_latent_eval/content_only_metrics_summary.csv)、[content_shuffle_summary.csv](outputs/m4_shared_latent_eval/content_shuffle_summary.csv)、[shifted_reconstruction_summary.csv](outputs/m4_shared_latent_eval/shifted_reconstruction_summary.csv) 和 [run_manifest.json](outputs/m4_shared_latent_eval/run_manifest.json)。
|
||||||
|
|
||||||
|
Fedora WSL 复现(复用已有五折 D5 检查点):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv run python -m q1.m4_shared_latent_eval --device cuda --seed 42 --probe-epochs 40 --decoder-epochs 40 --shuffle-repeats 20
|
||||||
|
```
|
||||||
|
|
||||||
|
### TSFA:冻结时间候选范围与局部内容选择
|
||||||
|
|
||||||
|
在上述 D5/M4 结果之上,本轮验证“先用时间确定候选范围,再在范围内按内容选择”能否同时改善时间定位和跨模态内容对应。五折 `M4_sourceTime` 检查点保持冻结;每个共享槽的 M4 注意力期望时间作为候选中心,主方案在归一化时间 `±0.10` 的范围内,用文本内容查询音频和视觉内容。局部语义分支只用训练折的同槽正例和相隔 `±2、±3、±5` 槽的难负例训练 40 个 epoch。它的 Q/K 不显式接收源时间码、位置编码或槽编号;M4 选中的内容仍可能间接携带时间。情绪标签没有进入训练。原有 BERT 文本、eGeMAPS 音频和 DeiT 视觉特征均保持不变。
|
||||||
|
|
||||||
|
比较包括 M3 无/有时间码、冻结 M4、有时间候选的 TSFA 主方案、乘以 M4 时间注意力的变体、20 次同宽度随机候选窗口,以及允许搜索全部源位置的全局变体。全部方法复用相同的 5 个 `video_id` 分组折、训练折归一化、每折相同的评估探针随机种子。100 条样本各作一次留出预测,涉及 37 个原视频。表中 AUC 是 Text–Audio、Text–Vision、Audio–Vision 三组“同槽对错槽”AUC 的逐样本平均,再按 `video_id` 宏平均;`0.5` 为机会水平。时间 MAE 是三个正向跨模态映射在归一化视频时间上的平均,越低越好。
|
||||||
|
|
||||||
|
| 方法 | 模型输出 AUC ↑ | 统一原始特征 AUC ↑ | 时间 MAE ↓ |
|
||||||
|
| --- | ---: | ---: | ---: |
|
||||||
|
| M3 无时间码 | 0.708 | 0.592 | 0.251 |
|
||||||
|
| M3 有时间码 | 0.565 | 0.516 | 0.099 |
|
||||||
|
| M4 有时间码 | 0.544 | 0.517 | **0.064** |
|
||||||
|
| TSFA 主方案 | 0.627 | 0.553 | 0.070 |
|
||||||
|
| TSFA 乘时间权重 | 0.627 | 0.551 | 0.068 |
|
||||||
|
| TSFA 随机窗口 | 0.632 | 0.571 | 0.251 |
|
||||||
|
| TSFA 全局候选 | **0.685** | **0.608** | 0.243 |
|
||||||
|
|
||||||
|
时间轴诊断进一步说明候选范围的作用。MVR 使用相邻槽期望时间倒退超过 `0.02` 的比例;跨度是末槽减首槽的归一化时间。熵只作分布诊断,不能单独排名。
|
||||||
|
|
||||||
|
| 方法 | Audio MVR ↓ | Vision MVR ↓ | Audio 跨度 ↑ | Vision 跨度 ↑ | Audio 熵 | Vision 熵 |
|
||||||
|
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
|
||||||
|
| M4 有时间码 | 0.0000 | 0.0000 | 0.816 | 0.822 | 0.846 | 0.720 |
|
||||||
|
| TSFA 主方案 | 0.0025 | 0.0004 | 0.796 | 0.821 | 0.538 | 0.510 |
|
||||||
|
| TSFA 随机窗口 | 0.4650 | 0.4745 | -0.017 | 0.010 | 0.533 | 0.500 |
|
||||||
|
| TSFA 全局候选 | 0.1406 | 0.0645 | 0.011 | 0.008 | 0.701 | 0.902 |
|
||||||
|
|
||||||
|
主方案的注意力更集中,且仍覆盖约 `99.9%` 的有效源位置;因此时间 MAE 的对照没有靠删除大批难对齐源位置取得优势。随机窗口同样能给出较低的熵,却有严重倒序和近零跨度,再次说明熵低不等于对齐正确。
|
||||||
|
|
||||||
|
“模型输出”沿用各方法自己的 Value/输出投影;“统一原始特征”把同一折中归一化后的 BERT、音频和 DeiT 源特征分别乘以各方法的对齐矩阵,再训练相同的线性对应探针。后者隔离了“选中哪些源位置”的贡献,更适合判断对齐矩阵本身。两者都只用训练视频拟合探针,留出视频打分。M3 无时间码的模型输出 Text–Audio AUC 达 `0.956`,统一原始特征后为 `0.725`;这说明原生注意力输出的查询相关变换也对它的高分有贡献,不能把原生输出 AUC 全部解释为正确的时间选择。
|
||||||
|
|
||||||
|
成对比较同一留出样本,并按 `video_id` 做 2,000 次 bootstrap:TSFA 主方案相对 M4 的统一原始特征 AUC 增加 `0.035`,95% 区间 `[0.023, 0.047]`;时间 MAE 增加 `0.0059`,区间 `[0.0053, 0.0065]`,即时间精度略有下降。它的统一原始特征 Text–Audio AUC 为 `0.639`,M4 为 `0.539`;Text–Vision `0.507`、Audio–Vision `0.511` 仍接近机会水平。原生输出的 Text–Audio AUC 为 `0.844`,M4 为 `0.616`;独立打乱留出片段内各模态槽顺序后,TSFA Text–Audio AUC 降至 `0.500`,说明该探针确实依赖槽顺序。
|
||||||
|
|
||||||
|
随机窗口相对主方案的统一原始特征 AUC 高 `0.018`,全局候选高 `0.055`;同时二者的时间 MAE 分别升至 `0.251` 和 `0.243`。随机窗口使用相同 `δ`,但边界截断使 Audio 平均候选比例为 `19.3%`,主方案为 `20.3%`;这一对照并未做到逐槽候选数量完全相等。因此局部 M4 候选范围的作用主要是**约束时间位置**,现有实验没有证明它能提高内容 AUC。乘时间权重比硬候选主方案把时间 MAE 从 `0.070` 降至 `0.068`,但统一原始特征 AUC 从 `0.553` 降至 `0.551`。这是一条可描述的时间—内容权衡,不宜合成为一个总分。
|
||||||
|
|
||||||
|
错位重构也只给出有限支持:错开 10 槽时,TSFA 的 Vision 内容重构误差增加 `0.00148`,95% 区间 `[0.00050, 0.00240]`;Audio 和 Text 目标的对应区间都包含 0。重构对象是聚合特征,不是波形或像素。当前最稳妥的判断是:**TSFA 比冻结 M4 保留了更多可匹配的 Text–Audio 内容,且维持了大部分时间结构;它尚未证明可靠的三模态语义对齐。**训练折的同槽标签来自时间/槽约定,M4 的时间先验来自时间戳,均不是人工事件真值;直接的事件 IoU/MATE 仍需人工标注。此轮只有一个对齐训练种子,视频组区间不包含训练种子波动。基于这轮七方法结果,暂不把 `δ`、RoPE 或 latent 长度消融写成已完成实验。
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
完整的逐片段指标、训练记录和探针检查点位于 [outputs/tsfa](outputs/tsfa/);可分享的汇总与图在 [report_bundle](outputs/tsfa/report_bundle/)。重点文件为 [七方法主表](outputs/tsfa/ablation_summary.csv)、[统一原始特征主表](outputs/tsfa/alignment_only_ablation_summary.csv)、[统一特征成对差值](outputs/tsfa/alignment_only_paired_contrasts.csv)、[三组配对和双向检索](outputs/tsfa/alignment_only_content_summary.csv)、[MVR/熵/覆盖率汇总](outputs/tsfa/tsfa_temporal_diagnostics_summary.csv)、[随机窗口统计](outputs/tsfa/candidate_window_stats.csv)、[运行清单](outputs/tsfa/run_manifest.json) 与 [统一特征运行清单](outputs/tsfa/alignment_only_run_manifest.json)。
|
||||||
|
|
||||||
|
Fedora WSL 复现(在 `deep_learning/Q1` 中执行,Python 环境由 `uv` 管理):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
uv run python -m q1.tsfa_experiment --device cuda --seed 42
|
||||||
|
uv run python -m q1.tsfa_alignment_only_eval --device cuda --seed 42
|
||||||
|
```
|
||||||
Binary file not shown.
@@ -0,0 +1,26 @@
|
|||||||
|
[project]
|
||||||
|
name = "q1"
|
||||||
|
version = "0.1.0"
|
||||||
|
requires-python = ">=3.14"
|
||||||
|
dependencies = [
|
||||||
|
"matplotlib>=3.11.2",
|
||||||
|
"mediapipe>=1.0.1",
|
||||||
|
"numpy>=2.5.2",
|
||||||
|
"opencv-python-headless>=5.0.0.93",
|
||||||
|
"openpyxl>=3.1.5",
|
||||||
|
"opensmile>=2.6.0",
|
||||||
|
"pillow>=12.3.0",
|
||||||
|
"scikit-learn>=1.9.1",
|
||||||
|
"soundfile>=0.14.0",
|
||||||
|
"torch>=2.14.0",
|
||||||
|
"torchvision>=0.29.0",
|
||||||
|
"transformers>=5.17.0",
|
||||||
|
]
|
||||||
|
|
||||||
|
[tool.uv.sources]
|
||||||
|
torch = { index = "pytorch" }
|
||||||
|
|
||||||
|
[[tool.uv.index]]
|
||||||
|
name = "pytorch"
|
||||||
|
url = "https://download.pytorch.org/whl/cu130"
|
||||||
|
explicit = true
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
"""Core data structures and alignment methods for Q1."""
|
||||||
|
|
||||||
|
from .alignment import align_fixed_windows, align_forced_timestamps
|
||||||
|
from .models import SharedLatentTimeline, TextAnchoredCrossAttention
|
||||||
|
from .types import AlignmentOutput, SequenceBatch
|
||||||
|
|
||||||
|
__all__ = [
|
||||||
|
"AlignmentOutput",
|
||||||
|
"SequenceBatch",
|
||||||
|
"SharedLatentTimeline",
|
||||||
|
"TextAnchoredCrossAttention",
|
||||||
|
"align_fixed_windows",
|
||||||
|
"align_forced_timestamps",
|
||||||
|
]
|
||||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,188 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from collections.abc import Sequence
|
||||||
|
|
||||||
|
import torch
|
||||||
|
from torch import Tensor
|
||||||
|
|
||||||
|
from .types import AlignmentOutput, MODALITIES, SequenceBatch
|
||||||
|
|
||||||
|
|
||||||
|
def uniform_time_intervals(durations: Tensor, grid_size: int) -> Tensor:
|
||||||
|
"""Return ``[B, K, 2]`` equal-duration windows in seconds."""
|
||||||
|
if durations.ndim != 1 or grid_size < 1:
|
||||||
|
raise ValueError("durations must be [B] and grid_size must be positive")
|
||||||
|
if bool((durations <= 0).any()) or not bool(torch.isfinite(durations).all()):
|
||||||
|
raise ValueError("durations must be finite and positive")
|
||||||
|
edges = torch.linspace(
|
||||||
|
0.0, 1.0, grid_size + 1, device=durations.device, dtype=durations.dtype
|
||||||
|
)[None, :] * durations[:, None]
|
||||||
|
return torch.stack((edges[:, :-1], edges[:, 1:]), dim=-1)
|
||||||
|
|
||||||
|
|
||||||
|
def word_intervals_to_grid(word_intervals: Tensor, grid_size: int) -> Tensor:
|
||||||
|
"""Resample ordered word spans into K consecutive text-order intervals.
|
||||||
|
|
||||||
|
Each grid slot covers an equal share of transcript word order. Its time
|
||||||
|
boundaries are interpolated from forced-alignment word boundaries, so long
|
||||||
|
and short words retain their actual duration on the audio/video timeline.
|
||||||
|
"""
|
||||||
|
if word_intervals.ndim != 2 or word_intervals.shape[1] != 2:
|
||||||
|
raise ValueError("word_intervals must have shape [word_count, 2]")
|
||||||
|
if word_intervals.shape[0] == 0 or grid_size < 1:
|
||||||
|
raise ValueError("at least one word interval and a positive grid size are required")
|
||||||
|
intervals = word_intervals.to(dtype=torch.float32)
|
||||||
|
if not bool(torch.isfinite(intervals).all()):
|
||||||
|
raise ValueError("word intervals must be finite")
|
||||||
|
if bool((intervals[:, 1] < intervals[:, 0]).any()):
|
||||||
|
raise ValueError("word interval end must not precede its start")
|
||||||
|
if bool((intervals[1:, 0] < intervals[:-1, 0]).any()):
|
||||||
|
raise ValueError("word intervals must be ordered by start time")
|
||||||
|
|
||||||
|
word_count = intervals.shape[0]
|
||||||
|
if word_count == 1:
|
||||||
|
boundaries = torch.cat((intervals[:1, 0], intervals[:1, 1]))
|
||||||
|
else:
|
||||||
|
between = (intervals[:-1, 1] + intervals[1:, 0]) / 2
|
||||||
|
boundaries = torch.cat((intervals[:1, 0], between, intervals[-1:, 1]))
|
||||||
|
boundaries = torch.cummax(boundaries, dim=0).values
|
||||||
|
|
||||||
|
positions = torch.linspace(
|
||||||
|
0, word_count, grid_size + 1, device=intervals.device, dtype=intervals.dtype
|
||||||
|
)
|
||||||
|
left = positions.floor().long().clamp(max=word_count)
|
||||||
|
right = (left + 1).clamp(max=word_count)
|
||||||
|
fraction = (positions - left.to(positions.dtype)).unsqueeze(-1)
|
||||||
|
time_edges = boundaries[left] + fraction.squeeze(-1) * (boundaries[right] - boundaries[left])
|
||||||
|
return torch.stack((time_edges[:-1], time_edges[1:]), dim=-1)
|
||||||
|
|
||||||
|
|
||||||
|
def index_alignment(valid: Tensor, grid_size: int) -> tuple[Tensor, int]:
|
||||||
|
"""Map equal ranges of valid sequence order to K grid slots."""
|
||||||
|
if valid.ndim != 2 or valid.dtype != torch.bool:
|
||||||
|
raise ValueError("valid must be a boolean [B, L] tensor")
|
||||||
|
if grid_size < 1:
|
||||||
|
raise ValueError("grid_size must be positive")
|
||||||
|
batch_size, length = valid.shape
|
||||||
|
result = torch.zeros(batch_size, grid_size, length, device=valid.device, dtype=torch.float32)
|
||||||
|
fallback_count = 0
|
||||||
|
for batch_index in range(batch_size):
|
||||||
|
positions = torch.nonzero(valid[batch_index], as_tuple=False).flatten()
|
||||||
|
count = positions.numel()
|
||||||
|
if count == 0:
|
||||||
|
raise ValueError("each sample must contain a valid position")
|
||||||
|
for grid_index in range(grid_size):
|
||||||
|
start = (grid_index * count) // grid_size
|
||||||
|
end = ((grid_index + 1) * count) // grid_size
|
||||||
|
if start == end:
|
||||||
|
source_index = min(int((grid_index + 0.5) * count / grid_size), count - 1)
|
||||||
|
result[batch_index, grid_index, positions[source_index]] = 1.0
|
||||||
|
fallback_count += 1
|
||||||
|
else:
|
||||||
|
chosen = positions[start:end]
|
||||||
|
result[batch_index, grid_index, chosen] = 1.0 / chosen.numel()
|
||||||
|
return result, fallback_count
|
||||||
|
|
||||||
|
|
||||||
|
def interval_alignment(times: Tensor, valid: Tensor, intervals: Tensor) -> tuple[Tensor, int]:
|
||||||
|
"""Create row-normalized interval membership weights with nearest-time fallback."""
|
||||||
|
if times.ndim != 2 or valid.shape != times.shape or valid.dtype != torch.bool:
|
||||||
|
raise ValueError("times and valid must have matching [B, L] shapes")
|
||||||
|
if intervals.ndim != 3 or intervals.shape[0] != times.shape[0] or intervals.shape[2] != 2:
|
||||||
|
raise ValueError("intervals must have shape [B, K, 2]")
|
||||||
|
batch_size, length = times.shape
|
||||||
|
grid_size = intervals.shape[1]
|
||||||
|
result = torch.zeros(batch_size, grid_size, length, device=times.device, dtype=torch.float32)
|
||||||
|
fallback_count = 0
|
||||||
|
for batch_index in range(batch_size):
|
||||||
|
valid_positions = torch.nonzero(valid[batch_index], as_tuple=False).flatten()
|
||||||
|
valid_times = times[batch_index, valid_positions]
|
||||||
|
for grid_index in range(grid_size):
|
||||||
|
start, end = intervals[batch_index, grid_index]
|
||||||
|
is_last = grid_index == grid_size - 1
|
||||||
|
in_window = (valid_times >= start) & (
|
||||||
|
(valid_times <= end) if is_last else (valid_times < end)
|
||||||
|
)
|
||||||
|
chosen = valid_positions[in_window]
|
||||||
|
if chosen.numel() > 0:
|
||||||
|
result[batch_index, grid_index, chosen] = 1.0 / chosen.numel()
|
||||||
|
else:
|
||||||
|
center = (start + end) / 2
|
||||||
|
nearest = valid_positions[torch.argmin((valid_times - center).abs())]
|
||||||
|
result[batch_index, grid_index, nearest] = 1.0
|
||||||
|
fallback_count += 1
|
||||||
|
return result, fallback_count
|
||||||
|
|
||||||
|
|
||||||
|
def _output_from_weights(
|
||||||
|
sequences: dict[str, SequenceBatch],
|
||||||
|
weights: dict[str, Tensor],
|
||||||
|
fallbacks: dict[str, int],
|
||||||
|
) -> AlignmentOutput:
|
||||||
|
aligned = {name: torch.bmm(weights[name].to(sequences[name].features.dtype), sequences[name].features)
|
||||||
|
for name in MODALITIES}
|
||||||
|
output = AlignmentOutput(weights=weights, aligned=aligned, fallback_rows=fallbacks)
|
||||||
|
output.validate({name: sequences[name].valid for name in MODALITIES})
|
||||||
|
return output
|
||||||
|
|
||||||
|
|
||||||
|
def align_fixed_windows(
|
||||||
|
sequences: dict[str, SequenceBatch], durations: Tensor, grid_size: int = 50
|
||||||
|
) -> AlignmentOutput:
|
||||||
|
"""M2: average each modality inside the same K equal-duration windows."""
|
||||||
|
intervals = uniform_time_intervals(durations, grid_size)
|
||||||
|
weights: dict[str, Tensor] = {}
|
||||||
|
fallbacks: dict[str, int] = {}
|
||||||
|
for name in MODALITIES:
|
||||||
|
weights[name], fallbacks[name] = interval_alignment(
|
||||||
|
sequences[name].times, sequences[name].valid, intervals
|
||||||
|
)
|
||||||
|
return _output_from_weights(sequences, weights, fallbacks)
|
||||||
|
|
||||||
|
|
||||||
|
def align_forced_timestamps(
|
||||||
|
sequences: dict[str, SequenceBatch],
|
||||||
|
word_intervals: Sequence[Tensor],
|
||||||
|
grid_size: int = 50,
|
||||||
|
) -> AlignmentOutput:
|
||||||
|
"""M1: use forced word timestamps to define text-ordered grid intervals."""
|
||||||
|
batch_size = sequences["text"].features.shape[0]
|
||||||
|
if len(word_intervals) != batch_size:
|
||||||
|
raise ValueError("provide one ordered word-interval array per sample")
|
||||||
|
device = sequences["text"].features.device
|
||||||
|
intervals = torch.stack(
|
||||||
|
[word_intervals_to_grid(spans.to(device), grid_size) for spans in word_intervals], dim=0
|
||||||
|
)
|
||||||
|
weights: dict[str, Tensor] = {}
|
||||||
|
fallbacks: dict[str, int] = {}
|
||||||
|
for name in MODALITIES:
|
||||||
|
weights[name], fallbacks[name] = interval_alignment(
|
||||||
|
sequences[name].times, sequences[name].valid, intervals
|
||||||
|
)
|
||||||
|
return _output_from_weights(sequences, weights, fallbacks)
|
||||||
|
|
||||||
|
|
||||||
|
def make_block_mask(
|
||||||
|
batch_size: int,
|
||||||
|
grid_size: int,
|
||||||
|
ratio: float,
|
||||||
|
device: torch.device | str,
|
||||||
|
generator: torch.Generator | None = None,
|
||||||
|
) -> Tensor:
|
||||||
|
"""Sample one continuous masked interval per sequence on the common grid."""
|
||||||
|
if not 0 < ratio < 1:
|
||||||
|
raise ValueError("ratio must be between zero and one")
|
||||||
|
if batch_size < 1 or grid_size < 2:
|
||||||
|
raise ValueError("batch_size must be positive and grid_size at least two")
|
||||||
|
block_length = min(max(1, round(grid_size * ratio)), grid_size - 1)
|
||||||
|
mask = torch.zeros(batch_size, grid_size, dtype=torch.bool, device=device)
|
||||||
|
starts = torch.randint(
|
||||||
|
0,
|
||||||
|
grid_size - block_length + 1,
|
||||||
|
(batch_size,),
|
||||||
|
device=device,
|
||||||
|
generator=generator,
|
||||||
|
)
|
||||||
|
offsets = torch.arange(block_length, device=device)
|
||||||
|
mask[torch.arange(batch_size, device=device)[:, None], starts[:, None] + offsets] = True
|
||||||
|
return mask
|
||||||
@@ -0,0 +1,776 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
import math
|
||||||
|
import platform
|
||||||
|
import random
|
||||||
|
import time
|
||||||
|
from collections import defaultdict
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Mapping, Sequence
|
||||||
|
|
||||||
|
import matplotlib
|
||||||
|
|
||||||
|
matplotlib.use("Agg")
|
||||||
|
import matplotlib.pyplot as plt
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
import torch.nn.functional as F
|
||||||
|
from torch import Tensor, nn
|
||||||
|
|
||||||
|
from .alignment import make_block_mask
|
||||||
|
from .compare_methods import AlignmentReconstructor, _write_csv
|
||||||
|
from .experiment_data import (
|
||||||
|
FeatureSample,
|
||||||
|
FeatureStats,
|
||||||
|
collate_feature_samples,
|
||||||
|
fit_feature_stats,
|
||||||
|
load_feature_samples,
|
||||||
|
)
|
||||||
|
from .losses import (
|
||||||
|
cross_modal_contrastive_loss,
|
||||||
|
temporal_span_loss,
|
||||||
|
weak_temporal_band_loss,
|
||||||
|
)
|
||||||
|
from .metrics import (
|
||||||
|
alignment_trajectory,
|
||||||
|
attention_row_similarity,
|
||||||
|
normalized_attention_entropy,
|
||||||
|
monotonicity_violation_rate,
|
||||||
|
)
|
||||||
|
from .models import SharedLatentTimeline, TextAnchoredCrossAttention
|
||||||
|
from .types import MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
EXAMPLE_ID = "-tPCytz4rww/12"
|
||||||
|
SIGMA = 0.10
|
||||||
|
GRID_SIZE = 50
|
||||||
|
HIDDEN_SIZE = 128
|
||||||
|
HEADS = 4
|
||||||
|
SINGLE_SAMPLE_STEPS = 1000
|
||||||
|
FULL_DATA_STEPS = 500
|
||||||
|
BATCH_SIZE = 8
|
||||||
|
LEARNING_RATE = 1e-3
|
||||||
|
LOG_INTERVAL = 20
|
||||||
|
|
||||||
|
|
||||||
|
def _seed_everything(seed: int) -> None:
|
||||||
|
random.seed(seed)
|
||||||
|
np.random.seed(seed)
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
torch.cuda.manual_seed_all(seed)
|
||||||
|
torch.backends.cudnn.deterministic = True
|
||||||
|
torch.backends.cudnn.benchmark = False
|
||||||
|
|
||||||
|
|
||||||
|
def _model(
|
||||||
|
method: str,
|
||||||
|
dimensions: Mapping[str, int],
|
||||||
|
*,
|
||||||
|
absolute_position_encoding: bool = False,
|
||||||
|
) -> nn.Module:
|
||||||
|
if method == "M3":
|
||||||
|
return TextAnchoredCrossAttention(
|
||||||
|
dimensions,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
hidden_size=HIDDEN_SIZE,
|
||||||
|
heads=HEADS,
|
||||||
|
dropout=0.0,
|
||||||
|
)
|
||||||
|
if method == "M4":
|
||||||
|
return SharedLatentTimeline(
|
||||||
|
dimensions,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
hidden_size=HIDDEN_SIZE,
|
||||||
|
heads=HEADS,
|
||||||
|
dropout=0.0,
|
||||||
|
absolute_position_encoding=absolute_position_encoding,
|
||||||
|
)
|
||||||
|
raise ValueError(f"unknown method: {method}")
|
||||||
|
|
||||||
|
|
||||||
|
def _centers(
|
||||||
|
method: str,
|
||||||
|
output: Any,
|
||||||
|
sequences: Mapping[str, Any],
|
||||||
|
durations: Tensor,
|
||||||
|
) -> tuple[dict[str, Tensor], Tensor]:
|
||||||
|
batch_size = durations.shape[0]
|
||||||
|
if method == "M3":
|
||||||
|
text_centers = torch.bmm(
|
||||||
|
output.weights["text"], sequences["text"].times.unsqueeze(-1)
|
||||||
|
).squeeze(-1)
|
||||||
|
text_centers = text_centers / durations[:, None].clamp_min(1e-8)
|
||||||
|
return {"audio": text_centers, "vision": text_centers}, text_centers
|
||||||
|
centers = (
|
||||||
|
torch.arange(GRID_SIZE, dtype=durations.dtype, device=durations.device) + 0.5
|
||||||
|
) / GRID_SIZE
|
||||||
|
centers = centers.unsqueeze(0).expand(batch_size, -1)
|
||||||
|
return {name: centers for name in MODALITIES}, centers
|
||||||
|
|
||||||
|
|
||||||
|
def _gaussian_targets(
|
||||||
|
method: str,
|
||||||
|
output: Any,
|
||||||
|
sequences: Mapping[str, Any],
|
||||||
|
durations: Tensor,
|
||||||
|
) -> dict[str, Tensor]:
|
||||||
|
centers, _ = _centers(method, output, sequences, durations)
|
||||||
|
target_names = ("audio", "vision") if method == "M3" else MODALITIES
|
||||||
|
targets: dict[str, Tensor] = {}
|
||||||
|
for name in target_names:
|
||||||
|
times = sequences[name].times / durations[:, None].clamp_min(1e-8)
|
||||||
|
difference = (times[:, None, :] - centers[name][:, :, None]) / SIGMA
|
||||||
|
logits = -0.5 * difference.square()
|
||||||
|
logits = logits.masked_fill(~sequences[name].valid[:, None, :], -torch.inf)
|
||||||
|
targets[name] = torch.softmax(logits, dim=-1)
|
||||||
|
return targets
|
||||||
|
|
||||||
|
|
||||||
|
def _gaussian_alignment_kl(output: Any, targets: Mapping[str, Tensor]) -> Tensor:
|
||||||
|
losses = []
|
||||||
|
for name, target in targets.items():
|
||||||
|
predicted = output.weights[name].clamp_min(1e-8)
|
||||||
|
safe_target = target.clamp_min(1e-12)
|
||||||
|
row_kl = (safe_target * (safe_target.log() - predicted.log())).sum(dim=-1)
|
||||||
|
losses.append(row_kl.mean())
|
||||||
|
return torch.stack(losses).mean()
|
||||||
|
|
||||||
|
|
||||||
|
def _training_components(
|
||||||
|
experiment: str,
|
||||||
|
method: str,
|
||||||
|
output: Any,
|
||||||
|
sequences: Mapping[str, Any],
|
||||||
|
durations: Tensor,
|
||||||
|
*,
|
||||||
|
decoder: AlignmentReconstructor | None,
|
||||||
|
block_generator: torch.Generator,
|
||||||
|
) -> tuple[Tensor, dict[str, Tensor], tuple[str, ...]]:
|
||||||
|
centers, _ = _centers(method, output, sequences, durations)
|
||||||
|
coverage_modalities = ("audio", "vision") if method == "M3" else MODALITIES
|
||||||
|
span = temporal_span_loss(
|
||||||
|
output,
|
||||||
|
{name: sequences[name].times for name in MODALITIES},
|
||||||
|
durations,
|
||||||
|
minimum_span=0.7,
|
||||||
|
modalities=coverage_modalities,
|
||||||
|
)
|
||||||
|
band = weak_temporal_band_loss(
|
||||||
|
output,
|
||||||
|
{name: sequences[name].times for name in MODALITIES},
|
||||||
|
durations,
|
||||||
|
centers,
|
||||||
|
margin=0.1,
|
||||||
|
)
|
||||||
|
targets = _gaussian_targets(method, output, sequences, durations)
|
||||||
|
align = _gaussian_alignment_kl(output, targets)
|
||||||
|
reconstruction = align.new_zeros(())
|
||||||
|
contrastive = align.new_zeros(())
|
||||||
|
|
||||||
|
if experiment == "D0":
|
||||||
|
total = 5.0 * span + 10.0 * band
|
||||||
|
active = ("span", "band", "total")
|
||||||
|
elif experiment in {"D1", "D2"}:
|
||||||
|
total = align
|
||||||
|
active = ("align", "total")
|
||||||
|
elif experiment == "D3":
|
||||||
|
if decoder is None:
|
||||||
|
raise ValueError("D3 requires the reconstruction decoder")
|
||||||
|
target_losses = []
|
||||||
|
batch_size, grid_size = output.aligned["text"].shape[:2]
|
||||||
|
for target in MODALITIES:
|
||||||
|
mask = make_block_mask(
|
||||||
|
batch_size,
|
||||||
|
grid_size,
|
||||||
|
0.2,
|
||||||
|
output.aligned[target].device,
|
||||||
|
generator=block_generator,
|
||||||
|
)
|
||||||
|
prediction = decoder(target, output.aligned, mask)
|
||||||
|
target_losses.append(
|
||||||
|
F.smooth_l1_loss(prediction[mask], output.aligned[target][mask])
|
||||||
|
)
|
||||||
|
reconstruction = torch.stack(target_losses).mean()
|
||||||
|
contrastive = cross_modal_contrastive_loss(output.aligned)
|
||||||
|
total = align + reconstruction + contrastive
|
||||||
|
active = ("align", "reconstruction", "contrastive", "total")
|
||||||
|
else:
|
||||||
|
raise ValueError(f"unknown experiment: {experiment}")
|
||||||
|
|
||||||
|
return total, {
|
||||||
|
"reconstruction": reconstruction,
|
||||||
|
"contrastive": contrastive,
|
||||||
|
"span": span,
|
||||||
|
"band": band,
|
||||||
|
"align": align,
|
||||||
|
"total": total,
|
||||||
|
}, active
|
||||||
|
|
||||||
|
|
||||||
|
def _attention_projection_parameters(model: nn.Module, method: str) -> tuple[list[nn.Parameter], nn.Parameter | None]:
|
||||||
|
if method == "M3":
|
||||||
|
layers = (model.audio_attention.attention, model.vision_attention.attention)
|
||||||
|
slot_parameter = None
|
||||||
|
else:
|
||||||
|
layers = tuple(model.attention[name].attention for name in MODALITIES)
|
||||||
|
slot_parameter = model.slots
|
||||||
|
parameters = [layer.in_proj_weight for layer in layers]
|
||||||
|
return parameters, slot_parameter
|
||||||
|
|
||||||
|
|
||||||
|
def _gradient_norms(
|
||||||
|
model: nn.Module,
|
||||||
|
method: str,
|
||||||
|
component_losses: Mapping[str, Tensor],
|
||||||
|
active_components: Sequence[str],
|
||||||
|
) -> dict[str, float | None]:
|
||||||
|
qk_parameters, slot_parameter = _attention_projection_parameters(model, method)
|
||||||
|
parameters = [*qk_parameters]
|
||||||
|
if slot_parameter is not None:
|
||||||
|
parameters.append(slot_parameter)
|
||||||
|
hidden_size = HIDDEN_SIZE
|
||||||
|
values: dict[str, float | None] = {}
|
||||||
|
for component in active_components:
|
||||||
|
loss = component_losses[component]
|
||||||
|
gradients = torch.autograd.grad(
|
||||||
|
loss,
|
||||||
|
parameters,
|
||||||
|
retain_graph=True,
|
||||||
|
allow_unused=True,
|
||||||
|
)
|
||||||
|
q_sq = torch.zeros((), device=loss.device)
|
||||||
|
k_sq = torch.zeros((), device=loss.device)
|
||||||
|
for gradient in gradients[: len(qk_parameters)]:
|
||||||
|
if gradient is None:
|
||||||
|
continue
|
||||||
|
q_sq = q_sq + gradient[:hidden_size].square().sum()
|
||||||
|
k_sq = k_sq + gradient[hidden_size : 2 * hidden_size].square().sum()
|
||||||
|
values[f"grad_{component}_WQ"] = float(q_sq.sqrt().item())
|
||||||
|
values[f"grad_{component}_WK"] = float(k_sq.sqrt().item())
|
||||||
|
if slot_parameter is not None:
|
||||||
|
slot_gradient = gradients[-1]
|
||||||
|
values[f"grad_{component}_Z"] = (
|
||||||
|
float(slot_gradient.norm().item()) if slot_gradient is not None else 0.0
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
values[f"grad_{component}_Z"] = None
|
||||||
|
return values
|
||||||
|
|
||||||
|
|
||||||
|
def _metric_rows(
|
||||||
|
method: str,
|
||||||
|
experiment: str,
|
||||||
|
sample: FeatureSample,
|
||||||
|
output: Any,
|
||||||
|
sequences: Mapping[str, Any],
|
||||||
|
durations: Tensor,
|
||||||
|
*,
|
||||||
|
stats: FeatureStats,
|
||||||
|
device: torch.device,
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
targets = _gaussian_targets(method, output, sequences, durations)
|
||||||
|
rows = []
|
||||||
|
for name in MODALITIES:
|
||||||
|
weights = output.weights[name]
|
||||||
|
valid = sequences[name].valid
|
||||||
|
trajectory = alignment_trajectory(weights, sequences[name].times, durations)
|
||||||
|
entropy = normalized_attention_entropy(weights, valid)
|
||||||
|
centers, _ = _centers(method, output, sequences, durations)
|
||||||
|
row: dict[str, Any] = {
|
||||||
|
"experiment": experiment,
|
||||||
|
"method": method,
|
||||||
|
"sample_id": sample.sample_id,
|
||||||
|
"modality": name,
|
||||||
|
"mvr": float(monotonicity_violation_rate(trajectory).mean().item()),
|
||||||
|
"normalized_entropy": float(entropy.mean().item()),
|
||||||
|
"c_row": float(attention_row_similarity(weights).mean().item()),
|
||||||
|
"c_far": float(attention_row_similarity(weights, min_separation=6).mean().item()),
|
||||||
|
"expected_time_start": float(trajectory[0, 0].item()),
|
||||||
|
"expected_time_end": float(trajectory[0, -1].item()),
|
||||||
|
"trajectory_span": float((trajectory[0, -1] - trajectory[0, 0]).item()),
|
||||||
|
"mean_absolute_time_center_error": float(
|
||||||
|
(trajectory - centers.get(name, trajectory.new_full(trajectory.shape, float("nan"))))
|
||||||
|
.abs()
|
||||||
|
.mean()
|
||||||
|
.item()
|
||||||
|
)
|
||||||
|
if name in centers
|
||||||
|
else None,
|
||||||
|
"gaussian_target_kl": None,
|
||||||
|
}
|
||||||
|
if name in targets:
|
||||||
|
target = targets[name].clamp_min(1e-12)
|
||||||
|
prediction = weights.clamp_min(1e-8)
|
||||||
|
row["gaussian_target_kl"] = float(
|
||||||
|
(target * (target.log() - prediction.log())).sum(dim=-1).mean().item()
|
||||||
|
)
|
||||||
|
rows.append(row)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def _example_arrays(sample: FeatureSample, output: Any, sequences: Mapping[str, Any], method: str) -> dict[str, np.ndarray]:
|
||||||
|
targets = _gaussian_targets(method, output, sequences, torch.tensor([sample.duration_s], device=output.weights["text"].device))
|
||||||
|
arrays: dict[str, np.ndarray] = {"sample_id": np.asarray(sample.sample_id)}
|
||||||
|
for name in MODALITIES:
|
||||||
|
length = len(sample.times[name])
|
||||||
|
arrays[f"weights_{name}"] = output.weights[name][0, :, :length].detach().cpu().numpy()
|
||||||
|
arrays[f"times_{name}_s"] = sample.times[name].astype(np.float32, copy=False)
|
||||||
|
arrays[f"valid_{name}"] = sample.valid[name]
|
||||||
|
trajectory = (
|
||||||
|
output.weights[name][0, :, :length]
|
||||||
|
@ torch.as_tensor(sample.times[name], dtype=torch.float32, device=output.weights[name].device)
|
||||||
|
/ sample.duration_s
|
||||||
|
)
|
||||||
|
arrays[f"trajectory_{name}"] = trajectory.detach().cpu().numpy()
|
||||||
|
if name in targets:
|
||||||
|
arrays[f"target_{name}"] = targets[name][0, :, :length].detach().cpu().numpy()
|
||||||
|
return arrays
|
||||||
|
|
||||||
|
|
||||||
|
def _plot_example(
|
||||||
|
path_prefix: Path,
|
||||||
|
sample: FeatureSample,
|
||||||
|
arrays: Mapping[str, np.ndarray],
|
||||||
|
method: str,
|
||||||
|
experiment: str,
|
||||||
|
) -> None:
|
||||||
|
modalities = ("audio", "vision") if method == "M3" else MODALITIES
|
||||||
|
fig, axes = plt.subplots(len(modalities), 2, figsize=(12, 4 * len(modalities)), constrained_layout=True)
|
||||||
|
if len(modalities) == 1:
|
||||||
|
axes = np.asarray([axes])
|
||||||
|
for row, name in enumerate(modalities):
|
||||||
|
times = arrays[f"times_{name}_s"] / max(sample.duration_s, 1e-8)
|
||||||
|
extent = (float(times[0]), float(times[-1]), 0.0, 1.0)
|
||||||
|
prediction = arrays[f"weights_{name}"]
|
||||||
|
image = axes[row, 0].imshow(
|
||||||
|
prediction,
|
||||||
|
origin="lower",
|
||||||
|
aspect="auto",
|
||||||
|
interpolation="nearest",
|
||||||
|
extent=extent,
|
||||||
|
cmap="magma",
|
||||||
|
)
|
||||||
|
axes[row, 0].set_title(f"{name}: learned A")
|
||||||
|
axes[row, 0].set_xlabel("source time / clip duration")
|
||||||
|
axes[row, 0].set_ylabel("slot index / K")
|
||||||
|
fig.colorbar(image, ax=axes[row, 0], fraction=0.046, pad=0.04)
|
||||||
|
target_key = f"target_{name}"
|
||||||
|
if target_key in arrays:
|
||||||
|
target = arrays[target_key]
|
||||||
|
image_target = axes[row, 1].imshow(
|
||||||
|
target,
|
||||||
|
origin="lower",
|
||||||
|
aspect="auto",
|
||||||
|
interpolation="nearest",
|
||||||
|
extent=extent,
|
||||||
|
cmap="magma",
|
||||||
|
)
|
||||||
|
axes[row, 1].set_title(f"{name}: Gaussian target P")
|
||||||
|
fig.colorbar(image_target, ax=axes[row, 1], fraction=0.046, pad=0.04)
|
||||||
|
else:
|
||||||
|
axes[row, 1].imshow(
|
||||||
|
np.zeros_like(prediction),
|
||||||
|
origin="lower",
|
||||||
|
aspect="auto",
|
||||||
|
extent=extent,
|
||||||
|
cmap="magma",
|
||||||
|
)
|
||||||
|
axes[row, 1].set_title(f"{name}: target not used by M3")
|
||||||
|
axes[row, 1].set_xlabel("source time / clip duration")
|
||||||
|
axes[row, 1].set_ylabel("slot index / K")
|
||||||
|
fig.suptitle(f"{experiment} {method} · {sample.sample_id}")
|
||||||
|
fig.savefig(path_prefix.with_name(path_prefix.name + "_heatmap.png"), dpi=160)
|
||||||
|
plt.close(fig)
|
||||||
|
|
||||||
|
fig, ax = plt.subplots(figsize=(8, 5), constrained_layout=True)
|
||||||
|
x = (np.arange(GRID_SIZE, dtype=np.float32) + 0.5) / GRID_SIZE
|
||||||
|
for name in MODALITIES:
|
||||||
|
trajectory = arrays[f"trajectory_{name}"]
|
||||||
|
ax.plot(x, trajectory, label=f"{name} actual")
|
||||||
|
if f"target_{name}" in arrays:
|
||||||
|
target_weights = arrays[f"target_{name}"]
|
||||||
|
target_time = target_weights @ arrays[f"times_{name}_s"] / max(sample.duration_s, 1e-8)
|
||||||
|
ax.plot(x, target_time, linestyle="--", alpha=0.7, label=f"{name} target")
|
||||||
|
ax.plot([0, 1], [0, 1], color="black", linestyle=":", alpha=0.6, label="uniform-time reference")
|
||||||
|
ax.set(xlim=(0, 1), ylim=(0, 1), xlabel="shared slot position", ylabel="expected normalized source time")
|
||||||
|
ax.grid(alpha=0.2)
|
||||||
|
ax.legend(fontsize=8, ncol=2)
|
||||||
|
ax.set_title(f"{experiment} {method} trajectory · {sample.sample_id}")
|
||||||
|
fig.savefig(path_prefix.with_name(path_prefix.name + "_trajectory.png"), dpi=160)
|
||||||
|
plt.close(fig)
|
||||||
|
|
||||||
|
|
||||||
|
def _batches(samples: Sequence[FeatureSample], batch_size: int, rng: np.random.Generator):
|
||||||
|
order = rng.permutation(len(samples)).tolist()
|
||||||
|
for start in range(0, len(order), batch_size):
|
||||||
|
yield [samples[index] for index in order[start : start + batch_size]]
|
||||||
|
|
||||||
|
|
||||||
|
def _evaluate_samples(
|
||||||
|
method: str,
|
||||||
|
experiment: str,
|
||||||
|
model: nn.Module,
|
||||||
|
samples: Sequence[FeatureSample],
|
||||||
|
stats: FeatureStats,
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
batch_size: int,
|
||||||
|
) -> tuple[list[dict[str, Any]], dict[str, np.ndarray] | None]:
|
||||||
|
rows: list[dict[str, Any]] = []
|
||||||
|
example_arrays: dict[str, np.ndarray] | None = None
|
||||||
|
model.eval()
|
||||||
|
rng = np.random.default_rng(0)
|
||||||
|
with torch.no_grad():
|
||||||
|
for batch_samples in _batches(samples, batch_size, rng):
|
||||||
|
sequences, durations, _ = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
output = model(sequences)
|
||||||
|
for index, sample in enumerate(batch_samples):
|
||||||
|
one_sequences = {
|
||||||
|
name: type(sequences[name])(
|
||||||
|
features=sequences[name].features[index : index + 1],
|
||||||
|
times=sequences[name].times[index : index + 1],
|
||||||
|
valid=sequences[name].valid[index : index + 1],
|
||||||
|
)
|
||||||
|
for name in MODALITIES
|
||||||
|
}
|
||||||
|
one_output = type(output)(
|
||||||
|
weights={name: output.weights[name][index : index + 1] for name in MODALITIES},
|
||||||
|
aligned={name: output.aligned[name][index : index + 1] for name in MODALITIES},
|
||||||
|
fallback_rows=output.fallback_rows,
|
||||||
|
)
|
||||||
|
one_duration = durations[index : index + 1]
|
||||||
|
rows.extend(
|
||||||
|
_metric_rows(
|
||||||
|
method,
|
||||||
|
experiment,
|
||||||
|
sample,
|
||||||
|
one_output,
|
||||||
|
one_sequences,
|
||||||
|
one_duration,
|
||||||
|
stats=stats,
|
||||||
|
device=device,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
if sample.sample_id == EXAMPLE_ID:
|
||||||
|
example_arrays = _example_arrays(sample, one_output, one_sequences, method)
|
||||||
|
return rows, example_arrays
|
||||||
|
|
||||||
|
|
||||||
|
def _run_trial(
|
||||||
|
experiment: str,
|
||||||
|
method: str,
|
||||||
|
tag: str,
|
||||||
|
train_samples: Sequence[FeatureSample],
|
||||||
|
eval_samples: Sequence[FeatureSample],
|
||||||
|
stats: FeatureStats,
|
||||||
|
output_dir: Path,
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
seed: int,
|
||||||
|
steps: int,
|
||||||
|
absolute_position_encoding: bool,
|
||||||
|
batch_size: int,
|
||||||
|
include_content_losses: bool,
|
||||||
|
) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
|
||||||
|
_seed_everything(seed)
|
||||||
|
dimensions = {name: train_samples[0].features[name].shape[1] for name in MODALITIES}
|
||||||
|
model = _model(method, dimensions, absolute_position_encoding=absolute_position_encoding).to(device)
|
||||||
|
decoder = AlignmentReconstructor(HIDDEN_SIZE, dropout=0.0).to(device) if include_content_losses else None
|
||||||
|
parameters = list(model.parameters()) + (list(decoder.parameters()) if decoder is not None else [])
|
||||||
|
optimizer = torch.optim.AdamW(parameters, lr=LEARNING_RATE, weight_decay=0.0)
|
||||||
|
block_generator = torch.Generator(device=device)
|
||||||
|
block_generator.manual_seed(seed + 73)
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
order = np.arange(len(train_samples))
|
||||||
|
cursor = 0
|
||||||
|
history: list[dict[str, Any]] = []
|
||||||
|
log_every = LOG_INTERVAL
|
||||||
|
active_names: tuple[str, ...] | None = None
|
||||||
|
|
||||||
|
print(
|
||||||
|
f"[{experiment} {tag}] samples={len(train_samples)} steps={steps} "
|
||||||
|
f"sinusoidal_PE={absolute_position_encoding} lr={LEARNING_RATE}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
model.train()
|
||||||
|
if decoder is not None:
|
||||||
|
decoder.train()
|
||||||
|
for step in range(1, steps + 1):
|
||||||
|
if len(train_samples) == 1:
|
||||||
|
batch_samples = [train_samples[0]]
|
||||||
|
else:
|
||||||
|
if cursor + batch_size > len(order):
|
||||||
|
order = rng.permutation(len(train_samples))
|
||||||
|
cursor = 0
|
||||||
|
indices = order[cursor : cursor + batch_size]
|
||||||
|
cursor += len(indices)
|
||||||
|
batch_samples = [train_samples[int(index)] for index in indices]
|
||||||
|
sequences, durations, _ = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
output = model(sequences)
|
||||||
|
total, losses, active_components = _training_components(
|
||||||
|
experiment,
|
||||||
|
method,
|
||||||
|
output,
|
||||||
|
sequences,
|
||||||
|
durations,
|
||||||
|
decoder=decoder,
|
||||||
|
block_generator=block_generator,
|
||||||
|
)
|
||||||
|
active_names = active_components
|
||||||
|
if not torch.isfinite(total):
|
||||||
|
raise FloatingPointError(f"non-finite {experiment}/{tag} loss at step {step}")
|
||||||
|
|
||||||
|
row: dict[str, Any] = {
|
||||||
|
"experiment": experiment,
|
||||||
|
"method": method,
|
||||||
|
"variant": tag,
|
||||||
|
"step": step,
|
||||||
|
"epoch": math.ceil(step / max(1, math.ceil(len(train_samples) / batch_size))),
|
||||||
|
"learning_rate": LEARNING_RATE,
|
||||||
|
"absolute_position_encoding": absolute_position_encoding,
|
||||||
|
"L_rec": float(losses["reconstruction"].detach().item()),
|
||||||
|
"L_con": float(losses["contrastive"].detach().item()),
|
||||||
|
"L_span": float(losses["span"].detach().item()),
|
||||||
|
"L_band": float(losses["band"].detach().item()),
|
||||||
|
"L_align": float(losses["align"].detach().item()),
|
||||||
|
"L_total": float(total.detach().item()),
|
||||||
|
}
|
||||||
|
if step == 1 or step % log_every == 0 or step == steps:
|
||||||
|
row.update(_gradient_norms(model, method, losses, active_components))
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
total.backward()
|
||||||
|
nn.utils.clip_grad_norm_(parameters, 2.0)
|
||||||
|
optimizer.step()
|
||||||
|
history.append(row)
|
||||||
|
if step == 1 or step % log_every == 0 or step == steps:
|
||||||
|
_write_csv(output_dir / f"{tag}_history.csv", history)
|
||||||
|
if step % 100 == 0 or step == steps:
|
||||||
|
diagnostic_component = "band" if experiment == "D0" else "align"
|
||||||
|
print(
|
||||||
|
f"[{experiment} {tag} step {step}/{steps}] "
|
||||||
|
f"total={row['L_total']:.4f} align={row['L_align']:.4f} "
|
||||||
|
f"span={row['L_span']:.4f} band={row['L_band']:.4f} "
|
||||||
|
f"grad-{diagnostic_component}(Q/K/Z)="
|
||||||
|
f"{row.get(f'grad_{diagnostic_component}_WQ', float('nan')):.3g}/"
|
||||||
|
f"{row.get(f'grad_{diagnostic_component}_WK', float('nan')):.3g}/"
|
||||||
|
f"{row.get(f'grad_{diagnostic_component}_Z', float('nan')) if row.get(f'grad_{diagnostic_component}_Z') is not None else 'NA'}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
_write_csv(output_dir / f"{tag}_history.csv", history)
|
||||||
|
checkpoint = output_dir / f"{tag}_checkpoint.pt"
|
||||||
|
torch.save(
|
||||||
|
{
|
||||||
|
"experiment": experiment,
|
||||||
|
"method": method,
|
||||||
|
"variant": tag,
|
||||||
|
"seed": seed,
|
||||||
|
"steps": steps,
|
||||||
|
"absolute_position_encoding": absolute_position_encoding,
|
||||||
|
"model_state_dict": model.state_dict(),
|
||||||
|
"decoder_state_dict": decoder.state_dict() if decoder is not None else None,
|
||||||
|
"history": history,
|
||||||
|
},
|
||||||
|
checkpoint,
|
||||||
|
)
|
||||||
|
metric_rows, example_arrays = _evaluate_samples(
|
||||||
|
method,
|
||||||
|
experiment,
|
||||||
|
model,
|
||||||
|
eval_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
batch_size=batch_size,
|
||||||
|
)
|
||||||
|
_write_csv(output_dir / f"{tag}_metrics.csv", metric_rows)
|
||||||
|
if example_arrays is not None:
|
||||||
|
np.savez_compressed(output_dir / f"{tag}_example_alignment.npz", **example_arrays)
|
||||||
|
example_sample = next(sample for sample in eval_samples if sample.sample_id == EXAMPLE_ID)
|
||||||
|
_plot_example(output_dir / tag, example_sample, example_arrays, method, experiment)
|
||||||
|
del model, decoder
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
return history, metric_rows
|
||||||
|
|
||||||
|
|
||||||
|
def run(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
started = time.time()
|
||||||
|
_seed_everything(args.seed)
|
||||||
|
if args.device == "auto":
|
||||||
|
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||||
|
else:
|
||||||
|
device = torch.device(args.device)
|
||||||
|
if device.type == "cuda" and not torch.cuda.is_available():
|
||||||
|
raise RuntimeError("CUDA was requested but is not available")
|
||||||
|
samples = load_feature_samples(args.feature_dir, args.manifest)
|
||||||
|
if len(samples) != 100:
|
||||||
|
raise ValueError(f"debug experiments expect the complete 100-sample set, found {len(samples)}")
|
||||||
|
sample_by_id = {sample.sample_id: sample for sample in samples}
|
||||||
|
if EXAMPLE_ID not in sample_by_id:
|
||||||
|
raise ValueError(f"required debug sample is absent: {EXAMPLE_ID}")
|
||||||
|
one_sample = sample_by_id[EXAMPLE_ID]
|
||||||
|
output_root = args.output_dir
|
||||||
|
output_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
full_stats = fit_feature_stats(samples)
|
||||||
|
single_stats = fit_feature_stats([one_sample])
|
||||||
|
all_metrics: list[dict[str, Any]] = []
|
||||||
|
all_history_summaries: list[dict[str, Any]] = []
|
||||||
|
|
||||||
|
specifications = [
|
||||||
|
("D0", "M3", "M3", [one_sample], [one_sample], single_stats, SINGLE_SAMPLE_STEPS, False, False),
|
||||||
|
("D0", "M4", "M4", [one_sample], [one_sample], single_stats, SINGLE_SAMPLE_STEPS, False, False),
|
||||||
|
("D1", "M3", "M3", [one_sample], [one_sample], single_stats, SINGLE_SAMPLE_STEPS, False, False),
|
||||||
|
("D1", "M4", "M4_noPE", [one_sample], [one_sample], single_stats, SINGLE_SAMPLE_STEPS, False, False),
|
||||||
|
("D1", "M4", "M4_sinPE", [one_sample], [one_sample], single_stats, SINGLE_SAMPLE_STEPS, True, False),
|
||||||
|
("D2", "M3", "M3", samples, samples, full_stats, FULL_DATA_STEPS, False, False),
|
||||||
|
("D2", "M4", "M4_sinPE", samples, samples, full_stats, FULL_DATA_STEPS, True, False),
|
||||||
|
("D3", "M3", "M3", samples, samples, full_stats, FULL_DATA_STEPS, False, True),
|
||||||
|
("D3", "M4", "M4_sinPE", samples, samples, full_stats, FULL_DATA_STEPS, True, True),
|
||||||
|
]
|
||||||
|
|
||||||
|
for experiment, method, tag, train_set, eval_set, stats, steps, use_pe, content_losses in specifications:
|
||||||
|
trial_dir = output_root / experiment
|
||||||
|
trial_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
history, metric_rows = _run_trial(
|
||||||
|
experiment,
|
||||||
|
method,
|
||||||
|
tag,
|
||||||
|
train_set,
|
||||||
|
eval_set,
|
||||||
|
stats,
|
||||||
|
trial_dir,
|
||||||
|
device=device,
|
||||||
|
seed=args.seed,
|
||||||
|
steps=steps,
|
||||||
|
absolute_position_encoding=use_pe,
|
||||||
|
batch_size=BATCH_SIZE,
|
||||||
|
include_content_losses=content_losses,
|
||||||
|
)
|
||||||
|
all_metrics.extend(metric_rows)
|
||||||
|
selected_steps = [
|
||||||
|
row
|
||||||
|
for row in history
|
||||||
|
if any(key.startswith("grad_") and value not in (None, "") for key, value in row.items())
|
||||||
|
]
|
||||||
|
summary: dict[str, Any] = {
|
||||||
|
"experiment": experiment,
|
||||||
|
"method": method,
|
||||||
|
"variant": tag,
|
||||||
|
"steps": steps,
|
||||||
|
"absolute_position_encoding": use_pe,
|
||||||
|
"final_L_total": history[-1]["L_total"],
|
||||||
|
"final_L_rec": history[-1]["L_rec"],
|
||||||
|
"final_L_con": history[-1]["L_con"],
|
||||||
|
"final_L_span": history[-1]["L_span"],
|
||||||
|
"final_L_band": history[-1]["L_band"],
|
||||||
|
"final_L_align": history[-1]["L_align"],
|
||||||
|
}
|
||||||
|
if selected_steps:
|
||||||
|
for component in ("span", "band", "align", "reconstruction", "contrastive", "total"):
|
||||||
|
for parameter in ("WQ", "WK", "Z"):
|
||||||
|
key = f"grad_{component}_{parameter}"
|
||||||
|
values = [float(row[key]) for row in selected_steps if row.get(key) not in (None, "")]
|
||||||
|
if values:
|
||||||
|
summary[f"mean_{key}"] = float(np.mean(values))
|
||||||
|
summary[f"final_{key}"] = values[-1]
|
||||||
|
summary["final_metric_rows"] = len(metric_rows)
|
||||||
|
all_history_summaries.append(summary)
|
||||||
|
print(
|
||||||
|
f"[done {experiment} {tag}] final total={summary['final_L_total']:.4f} "
|
||||||
|
f"align={summary['final_L_align']:.4f}; metrics={len(metric_rows)}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
_write_csv(output_root / "debug_summary.csv", all_history_summaries)
|
||||||
|
_write_csv(output_root / "per_sample_metrics.csv", all_metrics)
|
||||||
|
manifest = {
|
||||||
|
"created_utc": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"sample_count": len(samples),
|
||||||
|
"diagnostic_sample": EXAMPLE_ID,
|
||||||
|
"seed": args.seed,
|
||||||
|
"device": str(device),
|
||||||
|
"gpu_name": torch.cuda.get_device_name(device) if device.type == "cuda" else None,
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"experiments": [
|
||||||
|
{
|
||||||
|
"name": "D0",
|
||||||
|
"scope": "single sample; span + barycenter band only",
|
||||||
|
"steps_per_model": SINGLE_SAMPLE_STEPS,
|
||||||
|
"loss": "5 * L_span + 10 * L_band",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "D1",
|
||||||
|
"scope": "single sample; Gaussian target KL only",
|
||||||
|
"steps_per_model": SINGLE_SAMPLE_STEPS,
|
||||||
|
"sigma_normalized_time": SIGMA,
|
||||||
|
"M4_control": "no PE versus fixed sinusoidal PE",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "D2",
|
||||||
|
"scope": "all 100 clips; Gaussian target KL only; one seed; in-sample diagnostic",
|
||||||
|
"steps_per_model": FULL_DATA_STEPS,
|
||||||
|
"sigma_normalized_time": SIGMA,
|
||||||
|
"M4_position_encoding": "fixed sinusoidal",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "D3",
|
||||||
|
"scope": "all 100 clips; Gaussian KL + masked reconstruction + contrastive; one seed; in-sample diagnostic",
|
||||||
|
"steps_per_model": FULL_DATA_STEPS,
|
||||||
|
"sigma_normalized_time": SIGMA,
|
||||||
|
"M4_position_encoding": "fixed sinusoidal",
|
||||||
|
},
|
||||||
|
],
|
||||||
|
"optimizer": "AdamW",
|
||||||
|
"learning_rate": LEARNING_RATE,
|
||||||
|
"dropout": 0.0,
|
||||||
|
"batch_size": BATCH_SIZE,
|
||||||
|
"gradient_logging_interval_steps": LOG_INTERVAL,
|
||||||
|
"gradient_metrics": ["W_Q", "W_K", "M4 latent slots Z"],
|
||||||
|
"features_changed": False,
|
||||||
|
"M1_M2_changed": False,
|
||||||
|
"elapsed_seconds": time.time() - started,
|
||||||
|
"interpretation_limits": [
|
||||||
|
"D0 and D1 overfit one selected sample and diagnose optimization/representability only.",
|
||||||
|
"D2 and D3 train and evaluate on the same 100 clips; they diagnose whether the target can be optimized, not generalization.",
|
||||||
|
"Gaussian targets are weak temporal priors constructed from timestamps; they are not human alignment ground truth.",
|
||||||
|
"M4 absolute sinusoidal encoding is enabled only for D1's PE control and D2/D3; prior M1-M4 and v2 results are unchanged.",
|
||||||
|
],
|
||||||
|
}
|
||||||
|
(output_root / "run_manifest.json").write_text(
|
||||||
|
json.dumps(manifest, ensure_ascii=False, indent=2, allow_nan=False), encoding="utf-8"
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f"[all done] elapsed={manifest['elapsed_seconds']:.1f}s output={output_root}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
return manifest
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
project_dir = Path(__file__).resolve().parents[1]
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Diagnose gradient flow and temporal alignment learnability for M3/M4."
|
||||||
|
)
|
||||||
|
parser.add_argument("--feature-dir", type=Path, default=project_dir / "outputs/q1_features/features")
|
||||||
|
parser.add_argument("--manifest", type=Path, default=project_dir / "outputs/audit/manifest.csv")
|
||||||
|
parser.add_argument("--output-dir", type=Path, default=project_dir / "outputs/alignment_debug")
|
||||||
|
parser.add_argument("--device", default="auto", help="auto, cpu, or a torch device such as cuda:0")
|
||||||
|
parser.add_argument("--seed", type=int, default=42)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
run(args)
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,382 @@
|
|||||||
|
"""Grouped held-out check for learned source-time positional features."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
import platform
|
||||||
|
import time
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
from torch import nn
|
||||||
|
|
||||||
|
from .alignment_debug import (
|
||||||
|
EXAMPLE_ID,
|
||||||
|
GRID_SIZE,
|
||||||
|
HEADS,
|
||||||
|
HIDDEN_SIZE,
|
||||||
|
LEARNING_RATE,
|
||||||
|
_example_arrays,
|
||||||
|
_gaussian_alignment_kl,
|
||||||
|
_gaussian_targets,
|
||||||
|
_gradient_norms,
|
||||||
|
_metric_rows,
|
||||||
|
_plot_example,
|
||||||
|
_seed_everything,
|
||||||
|
)
|
||||||
|
from .experiment_data import (
|
||||||
|
FeatureSample,
|
||||||
|
FeatureStats,
|
||||||
|
collate_feature_samples,
|
||||||
|
fit_feature_stats,
|
||||||
|
load_feature_samples,
|
||||||
|
)
|
||||||
|
from .models import SharedLatentTimeline, TextAnchoredCrossAttention
|
||||||
|
from .types import MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
BATCH_SIZE = 8
|
||||||
|
DEFAULT_STEPS = 500
|
||||||
|
|
||||||
|
|
||||||
|
def _write_csv(path: Path, rows: list[dict[str, Any]]) -> None:
|
||||||
|
if not rows:
|
||||||
|
return
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
fields = list(dict.fromkeys(key for row in rows for key in row))
|
||||||
|
with path.open("w", newline="", encoding="utf-8-sig") as handle:
|
||||||
|
writer = csv.DictWriter(handle, fieldnames=fields)
|
||||||
|
writer.writeheader()
|
||||||
|
writer.writerows(rows)
|
||||||
|
|
||||||
|
|
||||||
|
def _make_model(method: str, dimensions: dict[str, int], source_time: bool) -> nn.Module:
|
||||||
|
if method == "M3":
|
||||||
|
return TextAnchoredCrossAttention(
|
||||||
|
dimensions,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
hidden_size=HIDDEN_SIZE,
|
||||||
|
heads=HEADS,
|
||||||
|
dropout=0.0,
|
||||||
|
source_time_encoding=source_time,
|
||||||
|
)
|
||||||
|
return SharedLatentTimeline(
|
||||||
|
dimensions,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
hidden_size=HIDDEN_SIZE,
|
||||||
|
heads=HEADS,
|
||||||
|
dropout=0.0,
|
||||||
|
absolute_position_encoding=True,
|
||||||
|
source_time_encoding=source_time,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _batches(samples: list[FeatureSample], rng: np.random.Generator):
|
||||||
|
order = rng.permutation(len(samples)).tolist()
|
||||||
|
for start in range(0, len(order), BATCH_SIZE):
|
||||||
|
yield [samples[index] for index in order[start : start + BATCH_SIZE]]
|
||||||
|
|
||||||
|
|
||||||
|
def _train_one(
|
||||||
|
*,
|
||||||
|
method: str,
|
||||||
|
variant: str,
|
||||||
|
source_time: bool,
|
||||||
|
train_samples: list[FeatureSample],
|
||||||
|
validation_samples: list[FeatureSample],
|
||||||
|
stats: FeatureStats,
|
||||||
|
example_id: str,
|
||||||
|
output_dir: Path,
|
||||||
|
device: torch.device,
|
||||||
|
steps: int,
|
||||||
|
seed: int,
|
||||||
|
) -> tuple[list[dict[str, Any]], list[dict[str, Any]], dict[str, Any]]:
|
||||||
|
_seed_everything(seed)
|
||||||
|
dimensions = {
|
||||||
|
name: train_samples[0].features[name].shape[1] for name in MODALITIES
|
||||||
|
}
|
||||||
|
model = _make_model(method, dimensions, source_time).to(device)
|
||||||
|
optimizer = torch.optim.AdamW(model.parameters(), lr=LEARNING_RATE, weight_decay=0.0)
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
history: list[dict[str, Any]] = []
|
||||||
|
print(
|
||||||
|
f"[D5 {variant}] train={len(train_samples)} heldout={len(validation_samples)} "
|
||||||
|
f"groups={len({s.group_id for s in train_samples})}/"
|
||||||
|
f"{len({s.group_id for s in validation_samples})} steps={steps}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
model.train()
|
||||||
|
for step in range(1, steps + 1):
|
||||||
|
batch_indices = rng.choice(
|
||||||
|
len(train_samples), size=min(BATCH_SIZE, len(train_samples)), replace=False
|
||||||
|
)
|
||||||
|
batch_samples = [train_samples[int(index)] for index in batch_indices]
|
||||||
|
sequences, durations, _ = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
output = model(sequences, durations) if source_time else model(sequences)
|
||||||
|
targets = _gaussian_targets(method, output, sequences, durations)
|
||||||
|
loss = _gaussian_alignment_kl(output, targets)
|
||||||
|
if not torch.isfinite(loss):
|
||||||
|
raise FloatingPointError(f"non-finite D5 loss for {variant} at step {step}")
|
||||||
|
row: dict[str, Any] = {
|
||||||
|
"experiment": "D5",
|
||||||
|
"method": method,
|
||||||
|
"variant": variant,
|
||||||
|
"step": step,
|
||||||
|
"L_align": float(loss.detach().item()),
|
||||||
|
"source_time_encoding": source_time,
|
||||||
|
}
|
||||||
|
if step == 1 or step % 20 == 0 or step == steps:
|
||||||
|
row.update(_gradient_norms(model, method, {"align": loss}, ("align",)))
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
loss.backward()
|
||||||
|
nn.utils.clip_grad_norm_(model.parameters(), 2.0)
|
||||||
|
optimizer.step()
|
||||||
|
history.append(row)
|
||||||
|
if step == 1 or step % 100 == 0 or step == steps:
|
||||||
|
grad_z = row.get("grad_align_Z")
|
||||||
|
grad_z_text = f"{grad_z:.3g}" if grad_z is not None else "NA"
|
||||||
|
print(
|
||||||
|
f"[D5 {variant} {step}/{steps}] KL={row['L_align']:.5f} "
|
||||||
|
f"grad_Q/K/Z={row.get('grad_align_WQ', 0):.3g}/"
|
||||||
|
f"{row.get('grad_align_WK', 0):.3g}/{grad_z_text}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
_write_csv(output_dir / "history.csv", history)
|
||||||
|
model.eval()
|
||||||
|
metric_rows: list[dict[str, Any]] = []
|
||||||
|
heldout_arrays: dict[str, np.ndarray] | None = None
|
||||||
|
with torch.no_grad():
|
||||||
|
eval_rng = np.random.default_rng(0)
|
||||||
|
for batch_samples in _batches(validation_samples, eval_rng):
|
||||||
|
sequences, durations, _ = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
output = model(sequences, durations) if source_time else model(sequences)
|
||||||
|
for index, sample in enumerate(batch_samples):
|
||||||
|
one_sequences = {
|
||||||
|
name: type(sequences[name])(
|
||||||
|
features=sequences[name].features[index : index + 1],
|
||||||
|
times=sequences[name].times[index : index + 1],
|
||||||
|
valid=sequences[name].valid[index : index + 1],
|
||||||
|
)
|
||||||
|
for name in MODALITIES
|
||||||
|
}
|
||||||
|
one_output = type(output)(
|
||||||
|
weights={name: output.weights[name][index : index + 1] for name in MODALITIES},
|
||||||
|
aligned={name: output.aligned[name][index : index + 1] for name in MODALITIES},
|
||||||
|
fallback_rows=output.fallback_rows,
|
||||||
|
)
|
||||||
|
rows = _metric_rows(
|
||||||
|
method,
|
||||||
|
"D5",
|
||||||
|
sample,
|
||||||
|
one_output,
|
||||||
|
one_sequences,
|
||||||
|
durations[index : index + 1],
|
||||||
|
stats=stats,
|
||||||
|
device=device,
|
||||||
|
)
|
||||||
|
for metric_row in rows:
|
||||||
|
metric_row["variant"] = variant
|
||||||
|
metric_row["source_time_encoding"] = source_time
|
||||||
|
metric_rows.extend(rows)
|
||||||
|
if sample.sample_id == example_id:
|
||||||
|
heldout_arrays = _example_arrays(
|
||||||
|
sample, one_output, one_sequences, method
|
||||||
|
)
|
||||||
|
|
||||||
|
_write_csv(output_dir / "heldout_metrics.csv", metric_rows)
|
||||||
|
if heldout_arrays is not None:
|
||||||
|
heldout_sample = next(s for s in validation_samples if s.sample_id == example_id)
|
||||||
|
np.savez_compressed(output_dir / "heldout_alignment.npz", **heldout_arrays)
|
||||||
|
_plot_example(output_dir / variant, heldout_sample, heldout_arrays, method, "D5 held-out")
|
||||||
|
checkpoint = output_dir / "checkpoint.pt"
|
||||||
|
torch.save(
|
||||||
|
{
|
||||||
|
"experiment": "D5",
|
||||||
|
"method": method,
|
||||||
|
"variant": variant,
|
||||||
|
"source_time_encoding": source_time,
|
||||||
|
"absolute_position_encoding": method == "M4",
|
||||||
|
"seed": seed,
|
||||||
|
"steps": steps,
|
||||||
|
"train_sample_ids": [sample.sample_id for sample in train_samples],
|
||||||
|
"validation_sample_ids": [sample.sample_id for sample in validation_samples],
|
||||||
|
"model_state_dict": model.state_dict(),
|
||||||
|
},
|
||||||
|
checkpoint,
|
||||||
|
)
|
||||||
|
training_summary = {
|
||||||
|
"experiment": "D5",
|
||||||
|
"method": method,
|
||||||
|
"variant": variant,
|
||||||
|
"source_time_encoding": source_time,
|
||||||
|
"final_training_kl": history[-1]["L_align"],
|
||||||
|
"heldout_sample_count": len(validation_samples),
|
||||||
|
"checkpoint": str(checkpoint),
|
||||||
|
}
|
||||||
|
del model
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
return metric_rows, history, training_summary
|
||||||
|
|
||||||
|
|
||||||
|
def run(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
started = time.time()
|
||||||
|
if args.device == "auto":
|
||||||
|
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||||
|
else:
|
||||||
|
device = torch.device(args.device)
|
||||||
|
if device.type == "cuda" and not torch.cuda.is_available():
|
||||||
|
raise RuntimeError("CUDA was requested but is unavailable")
|
||||||
|
samples = load_feature_samples(args.feature_dir, args.manifest)
|
||||||
|
by_id = {sample.sample_id: sample for sample in samples}
|
||||||
|
with args.splits.open("r", encoding="utf-8-sig") as handle:
|
||||||
|
folds = json.load(handle)
|
||||||
|
fold = next((item for item in folds if item["fold"] == args.fold), None)
|
||||||
|
if fold is None:
|
||||||
|
raise ValueError(f"fold {args.fold} is not present in {args.splits}")
|
||||||
|
train_samples = [by_id[sample_id] for sample_id in fold["train_sample_ids"]]
|
||||||
|
validation_samples = [by_id[sample_id] for sample_id in fold["validation_sample_ids"]]
|
||||||
|
train_groups = {sample.group_id for sample in train_samples}
|
||||||
|
validation_groups = {sample.group_id for sample in validation_samples}
|
||||||
|
if train_groups & validation_groups:
|
||||||
|
raise ValueError("train/validation video_id groups overlap")
|
||||||
|
example_id = args.example_id or validation_samples[0].sample_id
|
||||||
|
if example_id not in {sample.sample_id for sample in validation_samples}:
|
||||||
|
raise ValueError(f"held-out example is not in validation fold: {example_id}")
|
||||||
|
stats = fit_feature_stats(train_samples)
|
||||||
|
output_root = args.output_dir
|
||||||
|
output_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
specifications = (
|
||||||
|
("M3", "M3_noSourceTime", False),
|
||||||
|
("M3", "M3_sourceTime", True),
|
||||||
|
("M4", "M4_noSourceTime", False),
|
||||||
|
("M4", "M4_sourceTime", True),
|
||||||
|
)
|
||||||
|
all_metrics: list[dict[str, Any]] = []
|
||||||
|
all_history: list[dict[str, Any]] = []
|
||||||
|
summaries: list[dict[str, Any]] = []
|
||||||
|
for method, variant, source_time in specifications:
|
||||||
|
metrics, history, summary = _train_one(
|
||||||
|
method=method,
|
||||||
|
variant=variant,
|
||||||
|
source_time=source_time,
|
||||||
|
train_samples=train_samples,
|
||||||
|
validation_samples=validation_samples,
|
||||||
|
stats=stats,
|
||||||
|
example_id=example_id,
|
||||||
|
output_dir=output_root / variant,
|
||||||
|
device=device,
|
||||||
|
steps=args.steps,
|
||||||
|
seed=args.seed,
|
||||||
|
)
|
||||||
|
all_metrics.extend(metrics)
|
||||||
|
all_history.extend(history)
|
||||||
|
summaries.append(summary)
|
||||||
|
_write_csv(output_root / "per_sample_metrics.csv", all_metrics)
|
||||||
|
_write_csv(output_root / "training_history.csv", all_history)
|
||||||
|
|
||||||
|
aggregate_rows: list[dict[str, Any]] = []
|
||||||
|
metric_names = (
|
||||||
|
"mvr",
|
||||||
|
"normalized_entropy",
|
||||||
|
"c_row",
|
||||||
|
"trajectory_span",
|
||||||
|
"mean_absolute_time_center_error",
|
||||||
|
"gaussian_target_kl",
|
||||||
|
)
|
||||||
|
for variant, modality in sorted(
|
||||||
|
{(row["variant"], row["modality"]) for row in all_metrics}
|
||||||
|
):
|
||||||
|
rows = [row for row in all_metrics if row["variant"] == variant and row["modality"] == modality]
|
||||||
|
aggregate: dict[str, Any] = {
|
||||||
|
"variant": variant,
|
||||||
|
"modality": modality,
|
||||||
|
"sample_count": len(rows),
|
||||||
|
}
|
||||||
|
for name in metric_names:
|
||||||
|
values = [float(row[name]) for row in rows if row.get(name) not in (None, "")]
|
||||||
|
aggregate[f"mean_{name}"] = float(np.mean(values)) if values else ""
|
||||||
|
aggregate_rows.append(aggregate)
|
||||||
|
_write_csv(output_root / "heldout_summary.csv", aggregate_rows)
|
||||||
|
_write_csv(output_root / "training_summary.csv", summaries)
|
||||||
|
|
||||||
|
manifest = {
|
||||||
|
"created_utc": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"experiment": "D5",
|
||||||
|
"fold": args.fold,
|
||||||
|
"heldout_example": example_id,
|
||||||
|
"train_sample_count": len(train_samples),
|
||||||
|
"heldout_sample_count": len(validation_samples),
|
||||||
|
"train_video_ids": sorted(train_groups),
|
||||||
|
"heldout_video_ids": sorted(validation_groups),
|
||||||
|
"video_id_overlap": sorted(train_groups & validation_groups),
|
||||||
|
"seed": args.seed,
|
||||||
|
"device": str(device),
|
||||||
|
"gpu_name": torch.cuda.get_device_name(device) if device.type == "cuda" else None,
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"grid_size": GRID_SIZE,
|
||||||
|
"optimizer": "AdamW",
|
||||||
|
"learning_rate": LEARNING_RATE,
|
||||||
|
"steps_per_model": args.steps,
|
||||||
|
"batch_size": BATCH_SIZE,
|
||||||
|
"loss": "timestamp-derived Gaussian target KL only",
|
||||||
|
"source_position_encoding": "fixed Fourier time code added to source key only; value remains projected content",
|
||||||
|
"query_position_encoding": "M3 adds Fourier code at text-time centers; M4 uses fixed absolute sinusoidal slots plus the matching Fourier code at uniform slot centers",
|
||||||
|
"normalization_fit_on_train_only": True,
|
||||||
|
"variants": summaries,
|
||||||
|
"elapsed_seconds": time.time() - started,
|
||||||
|
"interpretation_limits": [
|
||||||
|
"This is one grouped video_id split and one seed; it is a focused held-out diagnostic, not a final method ranking.",
|
||||||
|
"The Gaussian timestamp target is a weak temporal prior, not human alignment ground truth.",
|
||||||
|
"The target supplies approximate time location; this experiment tests transfer of the time-conditioned attention mechanism, not semantic correctness by itself.",
|
||||||
|
],
|
||||||
|
}
|
||||||
|
(output_root / "run_manifest.json").write_text(
|
||||||
|
json.dumps(manifest, ensure_ascii=False, indent=2, allow_nan=False), encoding="utf-8"
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f"[D5 done] train={len(train_samples)} heldout={len(validation_samples)} "
|
||||||
|
f"video_id_groups={len(train_groups)}/{len(validation_groups)} "
|
||||||
|
f"elapsed={manifest['elapsed_seconds']:.1f}s output={output_root}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
return manifest
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
project = Path(__file__).resolve().parents[1]
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument("--steps", type=int, default=DEFAULT_STEPS)
|
||||||
|
parser.add_argument("--seed", type=int, default=42)
|
||||||
|
parser.add_argument("--fold", type=int, default=1)
|
||||||
|
parser.add_argument("--example-id", type=str, default=None)
|
||||||
|
parser.add_argument("--device", choices=("auto", "cpu", "cuda"), default="auto")
|
||||||
|
parser.add_argument(
|
||||||
|
"--feature-dir", type=Path, default=project / "outputs/q1_features/features"
|
||||||
|
)
|
||||||
|
parser.add_argument("--manifest", type=Path, default=project / "outputs/audit/manifest.csv")
|
||||||
|
parser.add_argument(
|
||||||
|
"--splits", type=Path, default=project / "outputs/method_comparison/splits.json"
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output-dir", type=Path, default=project / "outputs/alignment_debug/heldout"
|
||||||
|
)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
run(args)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,273 @@
|
|||||||
|
"""Single-clip test of explicit source-time identities in learned alignment."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
import platform
|
||||||
|
import time
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
from torch import nn
|
||||||
|
|
||||||
|
from .alignment_debug import (
|
||||||
|
EXAMPLE_ID,
|
||||||
|
GRID_SIZE,
|
||||||
|
HEADS,
|
||||||
|
HIDDEN_SIZE,
|
||||||
|
LEARNING_RATE,
|
||||||
|
SIGMA,
|
||||||
|
_example_arrays,
|
||||||
|
_gaussian_alignment_kl,
|
||||||
|
_gaussian_targets,
|
||||||
|
_gradient_norms,
|
||||||
|
_metric_rows,
|
||||||
|
_plot_example,
|
||||||
|
_seed_everything,
|
||||||
|
)
|
||||||
|
from .experiment_data import fit_feature_stats, load_feature_samples, collate_feature_samples
|
||||||
|
from .models import SharedLatentTimeline, TextAnchoredCrossAttention
|
||||||
|
from .types import MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
def _write_csv(path: Path, rows: list[dict[str, Any]]) -> None:
|
||||||
|
if not rows:
|
||||||
|
return
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
fields = list(dict.fromkeys(key for row in rows for key in row))
|
||||||
|
with path.open("w", newline="", encoding="utf-8-sig") as handle:
|
||||||
|
writer = csv.DictWriter(handle, fieldnames=fields)
|
||||||
|
writer.writeheader()
|
||||||
|
writer.writerows(rows)
|
||||||
|
|
||||||
|
|
||||||
|
def _make_model(method: str, dimensions: dict[str, int], source_time_encoding: bool) -> nn.Module:
|
||||||
|
if method == "M3":
|
||||||
|
return TextAnchoredCrossAttention(
|
||||||
|
dimensions,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
hidden_size=HIDDEN_SIZE,
|
||||||
|
heads=HEADS,
|
||||||
|
dropout=0.0,
|
||||||
|
source_time_encoding=source_time_encoding,
|
||||||
|
)
|
||||||
|
return SharedLatentTimeline(
|
||||||
|
dimensions,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
hidden_size=HIDDEN_SIZE,
|
||||||
|
heads=HEADS,
|
||||||
|
dropout=0.0,
|
||||||
|
absolute_position_encoding=True,
|
||||||
|
source_time_encoding=source_time_encoding,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _run_trial(
|
||||||
|
*,
|
||||||
|
method: str,
|
||||||
|
variant: str,
|
||||||
|
source_time_encoding: bool,
|
||||||
|
sample: Any,
|
||||||
|
stats: Any,
|
||||||
|
device: torch.device,
|
||||||
|
steps: int,
|
||||||
|
seed: int,
|
||||||
|
output_dir: Path,
|
||||||
|
) -> tuple[dict[str, Any], list[dict[str, Any]]]:
|
||||||
|
_seed_everything(seed)
|
||||||
|
dimensions = {name: sample.features[name].shape[1] for name in MODALITIES}
|
||||||
|
model = _make_model(method, dimensions, source_time_encoding).to(device)
|
||||||
|
optimizer = torch.optim.AdamW(model.parameters(), lr=LEARNING_RATE, weight_decay=0.0)
|
||||||
|
sequences, durations, _ = collate_feature_samples([sample], stats, device)
|
||||||
|
history: list[dict[str, Any]] = []
|
||||||
|
print(
|
||||||
|
f"[D4 {variant}] sample={sample.sample_id} steps={steps} "
|
||||||
|
f"source_time_encoding={source_time_encoding} device={device}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
model.train()
|
||||||
|
for step in range(1, steps + 1):
|
||||||
|
output = model(sequences, durations) if source_time_encoding else model(sequences)
|
||||||
|
targets = _gaussian_targets(method, output, sequences, durations)
|
||||||
|
loss = _gaussian_alignment_kl(output, targets)
|
||||||
|
if not torch.isfinite(loss):
|
||||||
|
raise FloatingPointError(f"non-finite D4 loss at {variant} step {step}")
|
||||||
|
row: dict[str, Any] = {
|
||||||
|
"experiment": "D4",
|
||||||
|
"method": method,
|
||||||
|
"variant": variant,
|
||||||
|
"step": step,
|
||||||
|
"L_align": float(loss.detach().item()),
|
||||||
|
"source_time_encoding": source_time_encoding,
|
||||||
|
}
|
||||||
|
if step == 1 or step % 20 == 0 or step == steps:
|
||||||
|
row.update(_gradient_norms(model, method, {"align": loss}, ("align",)))
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
loss.backward()
|
||||||
|
nn.utils.clip_grad_norm_(model.parameters(), 2.0)
|
||||||
|
optimizer.step()
|
||||||
|
history.append(row)
|
||||||
|
if step == 1 or step % 100 == 0 or step == steps:
|
||||||
|
grad_z = row.get("grad_align_Z")
|
||||||
|
grad_z_text = f"{grad_z:.3g}" if grad_z is not None else "NA"
|
||||||
|
print(
|
||||||
|
f"[D4 {variant} {step}/{steps}] KL={row['L_align']:.5f} "
|
||||||
|
f"grad_Q/K/Z={row.get('grad_align_WQ', 0):.3g}/"
|
||||||
|
f"{row.get('grad_align_WK', 0):.3g}/"
|
||||||
|
f"{grad_z_text}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
_write_csv(output_dir / "history.csv", history)
|
||||||
|
model.eval()
|
||||||
|
with torch.no_grad():
|
||||||
|
output = model(sequences, durations) if source_time_encoding else model(sequences)
|
||||||
|
metric_rows = _metric_rows(
|
||||||
|
method,
|
||||||
|
"D4",
|
||||||
|
sample,
|
||||||
|
output,
|
||||||
|
sequences,
|
||||||
|
durations,
|
||||||
|
stats=stats,
|
||||||
|
device=device,
|
||||||
|
)
|
||||||
|
arrays = _example_arrays(sample, output, sequences, method)
|
||||||
|
for row in metric_rows:
|
||||||
|
row["variant"] = variant
|
||||||
|
row["source_time_encoding"] = source_time_encoding
|
||||||
|
_write_csv(output_dir / "metrics.csv", metric_rows)
|
||||||
|
np.savez_compressed(output_dir / "example_alignment.npz", **arrays)
|
||||||
|
_plot_example(output_dir / variant, sample, arrays, method, "D4")
|
||||||
|
checkpoint = output_dir / "checkpoint.pt"
|
||||||
|
torch.save(
|
||||||
|
{
|
||||||
|
"experiment": "D4",
|
||||||
|
"method": method,
|
||||||
|
"variant": variant,
|
||||||
|
"source_time_encoding": source_time_encoding,
|
||||||
|
"absolute_position_encoding": method == "M4",
|
||||||
|
"seed": seed,
|
||||||
|
"steps": steps,
|
||||||
|
"model_state_dict": model.state_dict(),
|
||||||
|
},
|
||||||
|
checkpoint,
|
||||||
|
)
|
||||||
|
summary = {
|
||||||
|
"experiment": "D4",
|
||||||
|
"method": method,
|
||||||
|
"variant": variant,
|
||||||
|
"source_time_encoding": source_time_encoding,
|
||||||
|
"final_training_kl": history[-1]["L_align"],
|
||||||
|
"checkpoint": str(checkpoint),
|
||||||
|
}
|
||||||
|
del model
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
return summary, metric_rows
|
||||||
|
|
||||||
|
|
||||||
|
def run(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
start = time.time()
|
||||||
|
if args.device == "auto":
|
||||||
|
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||||
|
else:
|
||||||
|
device = torch.device(args.device)
|
||||||
|
if device.type == "cuda" and not torch.cuda.is_available():
|
||||||
|
raise RuntimeError("CUDA was requested but is unavailable")
|
||||||
|
samples = load_feature_samples(args.feature_dir, args.manifest)
|
||||||
|
sample_map = {sample.sample_id: sample for sample in samples}
|
||||||
|
if EXAMPLE_ID not in sample_map:
|
||||||
|
raise ValueError(f"diagnostic sample is missing: {EXAMPLE_ID}")
|
||||||
|
sample = sample_map[EXAMPLE_ID]
|
||||||
|
stats = fit_feature_stats([sample])
|
||||||
|
output_root = args.output_dir
|
||||||
|
output_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
specifications = (
|
||||||
|
("M3", "M3_noSourceTime", False),
|
||||||
|
("M3", "M3_sourceTime", True),
|
||||||
|
("M4", "M4_noSourceTime", False),
|
||||||
|
("M4", "M4_sourceTime", True),
|
||||||
|
)
|
||||||
|
summaries: list[dict[str, Any]] = []
|
||||||
|
all_metrics: list[dict[str, Any]] = []
|
||||||
|
for method, variant, use_source_time in specifications:
|
||||||
|
summary, metrics = _run_trial(
|
||||||
|
method=method,
|
||||||
|
variant=variant,
|
||||||
|
source_time_encoding=use_source_time,
|
||||||
|
sample=sample,
|
||||||
|
stats=stats,
|
||||||
|
device=device,
|
||||||
|
steps=args.steps,
|
||||||
|
seed=args.seed,
|
||||||
|
output_dir=output_root / variant,
|
||||||
|
)
|
||||||
|
summaries.append(summary)
|
||||||
|
all_metrics.extend(metrics)
|
||||||
|
_write_csv(output_root / "summary.csv", summaries)
|
||||||
|
_write_csv(output_root / "per_sample_metrics.csv", all_metrics)
|
||||||
|
manifest = {
|
||||||
|
"created_utc": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"experiment": "D4",
|
||||||
|
"sample": sample.sample_id,
|
||||||
|
"sample_count": 1,
|
||||||
|
"seed": args.seed,
|
||||||
|
"device": str(device),
|
||||||
|
"gpu_name": torch.cuda.get_device_name(device) if device.type == "cuda" else None,
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"grid_size": GRID_SIZE,
|
||||||
|
"sigma_normalized_time": SIGMA,
|
||||||
|
"optimizer": "AdamW",
|
||||||
|
"learning_rate": LEARNING_RATE,
|
||||||
|
"steps_per_model": args.steps,
|
||||||
|
"source_position_encoding": "fixed Fourier time code added to source key only; value remains the projected real feature",
|
||||||
|
"query_position_encoding": "M3 adds Fourier code at forced text-time centers; M4 keeps its fixed absolute sinusoidal slot code and adds the matching Fourier code at uniform slot centers",
|
||||||
|
"modalities": ["audio", "vision"],
|
||||||
|
"variants": summaries,
|
||||||
|
"elapsed_seconds": time.time() - start,
|
||||||
|
"interpretation_limits": [
|
||||||
|
"This is a one-sample overfit diagnostic, not a held-out accuracy result.",
|
||||||
|
"The timestamp-derived Gaussian target is a weak temporal prior, not human alignment ground truth.",
|
||||||
|
"The D4 variants test source/query time identity only; content features and downstream outputs still require separate evaluation.",
|
||||||
|
],
|
||||||
|
}
|
||||||
|
(output_root / "run_manifest.json").write_text(
|
||||||
|
json.dumps(manifest, ensure_ascii=False, indent=2, allow_nan=False), encoding="utf-8"
|
||||||
|
)
|
||||||
|
print(f"[D4 done] elapsed={manifest['elapsed_seconds']:.1f}s output={output_root}", flush=True)
|
||||||
|
return manifest
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
project = Path(__file__).resolve().parents[1]
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument("--steps", type=int, default=1000)
|
||||||
|
parser.add_argument("--seed", type=int, default=42)
|
||||||
|
parser.add_argument("--device", choices=("auto", "cpu", "cuda"), default="auto")
|
||||||
|
parser.add_argument(
|
||||||
|
"--feature-dir", type=Path, default=project / "outputs/q1_features/features"
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--manifest", type=Path, default=project / "outputs/audit/manifest.csv"
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output-dir", type=Path, default=project / "outputs/alignment_debug/source_time"
|
||||||
|
)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
run(args)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,246 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
import subprocess
|
||||||
|
from collections import Counter
|
||||||
|
from fractions import Fraction
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from openpyxl import load_workbook
|
||||||
|
|
||||||
|
|
||||||
|
def _identifier(value: Any) -> str:
|
||||||
|
if value is None:
|
||||||
|
return ""
|
||||||
|
if isinstance(value, float) and value.is_integer():
|
||||||
|
return str(int(value))
|
||||||
|
return str(value).strip()
|
||||||
|
|
||||||
|
|
||||||
|
def _read_labels(path: Path) -> list[dict[str, Any]]:
|
||||||
|
workbook = load_workbook(path, read_only=True, data_only=True)
|
||||||
|
sheet = workbook.active
|
||||||
|
rows = sheet.iter_rows(values_only=True)
|
||||||
|
header = next(rows, None)
|
||||||
|
if header is None:
|
||||||
|
raise ValueError(f"empty workbook: {path}")
|
||||||
|
names = [str(value).strip().lower() if value is not None else "" for value in header]
|
||||||
|
required = ("video_id", "clip_id", "text", "label", "annotation")
|
||||||
|
missing = set(required) - set(names)
|
||||||
|
if missing:
|
||||||
|
raise ValueError(f"missing required label columns: {sorted(missing)}")
|
||||||
|
indexes = {name: names.index(name) for name in required}
|
||||||
|
records = []
|
||||||
|
for values in rows:
|
||||||
|
if not values or all(value is None for value in values):
|
||||||
|
continue
|
||||||
|
record = {name: values[index] if index < len(values) else None for name, index in indexes.items()}
|
||||||
|
record["video_id"] = _identifier(record["video_id"])
|
||||||
|
record["clip_id"] = _identifier(record["clip_id"])
|
||||||
|
record["text"] = "" if record["text"] is None else str(record["text"]).strip()
|
||||||
|
record["annotation"] = "" if record["annotation"] is None else str(record["annotation"]).strip()
|
||||||
|
if record["label"] is not None:
|
||||||
|
try:
|
||||||
|
record["label"] = float(record["label"])
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
pass
|
||||||
|
records.append(record)
|
||||||
|
workbook.close()
|
||||||
|
return records
|
||||||
|
|
||||||
|
|
||||||
|
def _probe_video(path: Path) -> dict[str, Any]:
|
||||||
|
result = subprocess.run(
|
||||||
|
[
|
||||||
|
"ffprobe", "-v", "error", "-show_entries",
|
||||||
|
"format=duration:stream=codec_type,codec_name,width,height,avg_frame_rate,r_frame_rate,nb_frames,sample_rate,channels:frame=media_type,best_effort_timestamp_time,pkt_duration_time",
|
||||||
|
"-show_frames", "-of", "json", str(path),
|
||||||
|
],
|
||||||
|
check=True,
|
||||||
|
capture_output=True,
|
||||||
|
text=True,
|
||||||
|
)
|
||||||
|
payload = json.loads(result.stdout)
|
||||||
|
streams = payload.get("streams", [])
|
||||||
|
video = next((stream for stream in streams if stream.get("codec_type") == "video"), {})
|
||||||
|
audio = next((stream for stream in streams if stream.get("codec_type") == "audio"), {})
|
||||||
|
fps_text = video.get("avg_frame_rate") or video.get("r_frame_rate") or "0/1"
|
||||||
|
try:
|
||||||
|
fps = float(Fraction(fps_text))
|
||||||
|
except (ValueError, ZeroDivisionError):
|
||||||
|
fps = 0.0
|
||||||
|
container_duration = float(payload.get("format", {}).get("duration", 0.0))
|
||||||
|
frame_times: dict[str, list[tuple[float, float]]] = {"video": [], "audio": []}
|
||||||
|
for frame in payload.get("frames", []):
|
||||||
|
media_type = frame.get("media_type")
|
||||||
|
timestamp = frame.get("best_effort_timestamp_time")
|
||||||
|
if media_type not in frame_times or timestamp is None:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
start = float(timestamp)
|
||||||
|
packet_duration = float(frame.get("pkt_duration_time", 0.0) or 0.0)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
continue
|
||||||
|
frame_times[media_type].append((start, packet_duration))
|
||||||
|
|
||||||
|
video_times = frame_times["video"]
|
||||||
|
audio_times = frame_times["audio"]
|
||||||
|
fallback_frame_duration = 1.0 / fps if fps > 0 else 0.0
|
||||||
|
video_start = min((item[0] for item in video_times), default=0.0)
|
||||||
|
video_end = max((start + (packet_duration or fallback_frame_duration) for start, packet_duration in video_times), default=container_duration)
|
||||||
|
audio_start = min((item[0] for item in audio_times), default=0.0)
|
||||||
|
audio_end = max((start + packet_duration for start, packet_duration in audio_times), default=container_duration)
|
||||||
|
decoded_duration = max(0.0, video_end - video_start)
|
||||||
|
audio_duration = max(0.0, audio_end - audio_start)
|
||||||
|
return {
|
||||||
|
"duration_s": decoded_duration or container_duration,
|
||||||
|
"container_duration_s": container_duration,
|
||||||
|
"audio_duration_s": audio_duration,
|
||||||
|
"video_timeline_start_s": video_start,
|
||||||
|
"video_timeline_end_s": video_end,
|
||||||
|
"audio_timeline_start_s": audio_start,
|
||||||
|
"audio_timeline_end_s": audio_end,
|
||||||
|
"video_codec": video.get("codec_name", ""),
|
||||||
|
"width": video.get("width", ""),
|
||||||
|
"height": video.get("height", ""),
|
||||||
|
"fps": fps,
|
||||||
|
"video_frames": len(video_times),
|
||||||
|
"container_video_frames": video.get("nb_frames", ""),
|
||||||
|
"audio_codec": audio.get("codec_name", ""),
|
||||||
|
"audio_sample_rate": audio.get("sample_rate", ""),
|
||||||
|
"audio_channels": audio.get("channels", ""),
|
||||||
|
"has_audio": bool(audio),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def audit_dataset(video_root: Path, label_file: Path, output_dir: Path, expected_count: int = 100) -> dict[str, Any]:
|
||||||
|
records = _read_labels(label_file)
|
||||||
|
videos: dict[tuple[str, str], Path] = {}
|
||||||
|
duplicate_video_files: list[str] = []
|
||||||
|
for video in sorted(video_root.rglob("*.mp4")):
|
||||||
|
key = (video.parent.name, video.stem)
|
||||||
|
if key in videos:
|
||||||
|
duplicate_video_files.append(str(video))
|
||||||
|
else:
|
||||||
|
videos[key] = video
|
||||||
|
|
||||||
|
keys = [(str(row["video_id"]), str(row["clip_id"])) for row in records]
|
||||||
|
counts = Counter(keys)
|
||||||
|
duplicate_rows = [list(key) for key, count in counts.items() if count > 1]
|
||||||
|
matched_keys = set(keys) & set(videos)
|
||||||
|
missing_keys = [key for key in keys if key not in videos]
|
||||||
|
extra_keys = sorted(set(videos) - set(keys))
|
||||||
|
class_mismatches = []
|
||||||
|
probe_errors: list[dict[str, str]] = []
|
||||||
|
probe_by_key: dict[tuple[str, str], dict[str, Any]] = {}
|
||||||
|
for key in matched_keys:
|
||||||
|
try:
|
||||||
|
probe_by_key[key] = _probe_video(videos[key])
|
||||||
|
except (OSError, subprocess.CalledProcessError, json.JSONDecodeError, ValueError) as error:
|
||||||
|
probe_errors.append({"video_id": key[0], "clip_id": key[1], "error": str(error)})
|
||||||
|
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
manifest_path = output_dir / "manifest.csv"
|
||||||
|
with manifest_path.open("w", encoding="utf-8-sig", newline="") as file:
|
||||||
|
writer = csv.DictWriter(
|
||||||
|
file,
|
||||||
|
fieldnames=(
|
||||||
|
"video_id", "clip_id", "group_id", "text", "label", "annotation", "video_path",
|
||||||
|
"video_exists", "duration_s", "container_duration_s", "audio_duration_s",
|
||||||
|
"video_timeline_start_s", "video_timeline_end_s", "audio_timeline_start_s",
|
||||||
|
"audio_timeline_end_s", "video_codec", "width", "height", "fps",
|
||||||
|
"video_frames", "container_video_frames", "audio_codec", "audio_sample_rate",
|
||||||
|
"audio_channels", "has_audio",
|
||||||
|
),
|
||||||
|
)
|
||||||
|
writer.writeheader()
|
||||||
|
for row, key in zip(records, keys):
|
||||||
|
label = row["label"]
|
||||||
|
annotation = str(row["annotation"]).strip().lower()
|
||||||
|
expected_class = "negative" if isinstance(label, (int, float)) and label < 0 else (
|
||||||
|
"positive" if isinstance(label, (int, float)) and label > 0 else "neutral"
|
||||||
|
)
|
||||||
|
if annotation in {"negative", "neutral", "positive"} and annotation != expected_class:
|
||||||
|
class_mismatches.append({"video_id": key[0], "clip_id": key[1], "label": label, "annotation": annotation})
|
||||||
|
path = videos.get(key)
|
||||||
|
probe = probe_by_key.get(key, {})
|
||||||
|
writer.writerow({
|
||||||
|
"video_id": key[0],
|
||||||
|
"clip_id": key[1],
|
||||||
|
"group_id": key[0],
|
||||||
|
"text": row["text"],
|
||||||
|
"label": row["label"],
|
||||||
|
"annotation": row["annotation"],
|
||||||
|
"video_path": str(path.relative_to(video_root)) if path else "",
|
||||||
|
"video_exists": bool(path),
|
||||||
|
**probe,
|
||||||
|
})
|
||||||
|
|
||||||
|
durations = [info["duration_s"] for info in probe_by_key.values() if info["duration_s"] > 0]
|
||||||
|
duration_out_of_range = [
|
||||||
|
{"video_id": key[0], "clip_id": key[1], "duration_s": info["duration_s"]}
|
||||||
|
for key, info in probe_by_key.items()
|
||||||
|
if not (2.648 <= info["duration_s"] <= 34.567)
|
||||||
|
]
|
||||||
|
audio_missing = [list(key) for key, info in probe_by_key.items() if not info["has_audio"]]
|
||||||
|
|
||||||
|
summary = {
|
||||||
|
"video_root": str(video_root),
|
||||||
|
"label_file": str(label_file),
|
||||||
|
"expected_count": expected_count,
|
||||||
|
"label_rows": len(records),
|
||||||
|
"unique_video_clip_pairs": len(set(keys)),
|
||||||
|
"unique_video_ids": len({key[0] for key in keys}),
|
||||||
|
"video_files": len(videos),
|
||||||
|
"matched_samples": len(matched_keys),
|
||||||
|
"coverage_rate": len(matched_keys) / max(len(records), 1),
|
||||||
|
"duration_min_s": min(durations) if durations else None,
|
||||||
|
"duration_max_s": max(durations) if durations else None,
|
||||||
|
"stated_duration_range_s": [2.648, 34.567],
|
||||||
|
"duration_out_of_range": duration_out_of_range,
|
||||||
|
"audio_stream_missing": audio_missing,
|
||||||
|
"media_probe_errors": probe_errors,
|
||||||
|
"missing_video_pairs": [list(key) for key in missing_keys],
|
||||||
|
"unlisted_video_pairs": [list(key) for key in extra_keys],
|
||||||
|
"duplicate_label_pairs": duplicate_rows,
|
||||||
|
"duplicate_video_files": duplicate_video_files,
|
||||||
|
"label_class_mismatches": class_mismatches,
|
||||||
|
"manifest": str(manifest_path),
|
||||||
|
}
|
||||||
|
summary["coverage_status"] = (
|
||||||
|
"complete_with_metadata_warnings" if duration_out_of_range else "complete"
|
||||||
|
)
|
||||||
|
summary_path = output_dir / "coverage_summary.json"
|
||||||
|
summary_path.write_text(json.dumps(summary, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||||
|
return summary
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
project_dir = Path(__file__).resolve().parents[1]
|
||||||
|
repo_dir = project_dir.parent
|
||||||
|
default_data = repo_dir / "E题数据" / "附件1-数据集原始多模态样本" / "MOSEI数据集部分原始视频-100条"
|
||||||
|
parser = argparse.ArgumentParser(description="Audit the 100 raw-video Q1 samples and export a manifest.")
|
||||||
|
parser.add_argument("--video-root", type=Path, default=default_data)
|
||||||
|
parser.add_argument("--labels", type=Path, default=default_data / "label-100.xlsx")
|
||||||
|
parser.add_argument("--output-dir", type=Path, default=project_dir / "outputs" / "audit")
|
||||||
|
parser.add_argument("--expected-count", type=int, default=100)
|
||||||
|
args = parser.parse_args()
|
||||||
|
summary = audit_dataset(args.video_root, args.labels, args.output_dir, args.expected_count)
|
||||||
|
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||||
|
complete = (
|
||||||
|
summary["label_rows"] == args.expected_count
|
||||||
|
and summary["unique_video_clip_pairs"] == args.expected_count
|
||||||
|
and summary["matched_samples"] == args.expected_count
|
||||||
|
and not summary["duplicate_label_pairs"]
|
||||||
|
and not summary["label_class_mismatches"]
|
||||||
|
and not summary["audio_stream_missing"]
|
||||||
|
and not summary["media_probe_errors"]
|
||||||
|
)
|
||||||
|
return 0 if complete else 1
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,945 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
import math
|
||||||
|
import platform
|
||||||
|
import random
|
||||||
|
import statistics
|
||||||
|
import time
|
||||||
|
from collections import defaultdict
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Mapping, Sequence
|
||||||
|
|
||||||
|
import matplotlib
|
||||||
|
|
||||||
|
matplotlib.use("Agg")
|
||||||
|
import matplotlib.pyplot as plt
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
from sklearn.model_selection import GroupKFold
|
||||||
|
from torch import Tensor, nn
|
||||||
|
|
||||||
|
from .alignment import align_fixed_windows, align_forced_timestamps, make_block_mask
|
||||||
|
from .experiment_data import (
|
||||||
|
FeatureSample,
|
||||||
|
FeatureStats,
|
||||||
|
collate_feature_samples,
|
||||||
|
fit_feature_stats,
|
||||||
|
load_feature_samples,
|
||||||
|
)
|
||||||
|
from .experiment_probes import (
|
||||||
|
RetrievalProjection,
|
||||||
|
run_frozen_emotion_probe,
|
||||||
|
run_reconstruction_probe,
|
||||||
|
run_retrieval_probe,
|
||||||
|
)
|
||||||
|
from .metrics import (
|
||||||
|
attention_row_similarity,
|
||||||
|
alignment_trajectory,
|
||||||
|
attention_width80,
|
||||||
|
monotonicity_violation_rate,
|
||||||
|
normalized_attention_entropy,
|
||||||
|
)
|
||||||
|
from .models import SharedLatentTimeline, TextAnchoredCrossAttention
|
||||||
|
from .types import AlignmentOutput, MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
class AlignmentReconstructor(nn.Module):
|
||||||
|
"""Shared M3/M4 training decoder: predict one stream from the other two."""
|
||||||
|
|
||||||
|
def __init__(self, hidden_size: int = 128, dropout: float = 0.1) -> None:
|
||||||
|
super().__init__()
|
||||||
|
self.decoders = nn.ModuleDict(
|
||||||
|
{
|
||||||
|
target: nn.Sequential(
|
||||||
|
nn.Linear(hidden_size * 2, hidden_size),
|
||||||
|
nn.GELU(),
|
||||||
|
nn.Dropout(dropout),
|
||||||
|
nn.Linear(hidden_size, hidden_size),
|
||||||
|
)
|
||||||
|
for target in MODALITIES
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
def forward(self, target: str, sources: Mapping[str, Tensor], mask: Tensor) -> Tensor:
|
||||||
|
values = [
|
||||||
|
sources[name].masked_fill(mask.unsqueeze(-1), 0.0)
|
||||||
|
for name in MODALITIES
|
||||||
|
if name != target
|
||||||
|
]
|
||||||
|
return self.decoders[target](torch.cat(values, dim=-1))
|
||||||
|
|
||||||
|
|
||||||
|
def _seed_everything(seed: int) -> None:
|
||||||
|
random.seed(seed)
|
||||||
|
np.random.seed(seed)
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
torch.cuda.manual_seed_all(seed)
|
||||||
|
if hasattr(torch.backends, "cudnn"):
|
||||||
|
torch.backends.cudnn.deterministic = True
|
||||||
|
torch.backends.cudnn.benchmark = False
|
||||||
|
|
||||||
|
|
||||||
|
def _batches(
|
||||||
|
samples: Sequence[FeatureSample], batch_size: int, *, shuffle: bool, rng: np.random.Generator
|
||||||
|
) -> list[list[FeatureSample]]:
|
||||||
|
if shuffle:
|
||||||
|
order = rng.permutation(len(samples)).tolist()
|
||||||
|
else:
|
||||||
|
order = list(range(len(samples)))
|
||||||
|
return [[samples[index] for index in order[start : start + batch_size]]
|
||||||
|
for start in range(0, len(order), batch_size)]
|
||||||
|
|
||||||
|
|
||||||
|
def _training_objective(
|
||||||
|
output: AlignmentOutput,
|
||||||
|
sequences: Mapping[str, Any],
|
||||||
|
durations: Tensor,
|
||||||
|
decoder: AlignmentReconstructor,
|
||||||
|
generator: torch.Generator,
|
||||||
|
*,
|
||||||
|
method: str = "M3",
|
||||||
|
loss_variant: str = "v1",
|
||||||
|
) -> tuple[Tensor, dict[str, Tensor]]:
|
||||||
|
reconstruction_terms = []
|
||||||
|
batch_size, grid_size = output.aligned["text"].shape[:2]
|
||||||
|
for target in MODALITIES:
|
||||||
|
mask = make_block_mask(
|
||||||
|
batch_size,
|
||||||
|
grid_size,
|
||||||
|
0.2,
|
||||||
|
output.aligned[target].device,
|
||||||
|
generator=generator,
|
||||||
|
)
|
||||||
|
prediction = decoder(target, output.aligned, mask)
|
||||||
|
reconstruction_terms.append(
|
||||||
|
nn.functional.smooth_l1_loss(prediction[mask], output.aligned[target][mask])
|
||||||
|
)
|
||||||
|
reconstruction = torch.stack(reconstruction_terms).mean()
|
||||||
|
times = {name: sequences[name].times for name in MODALITIES}
|
||||||
|
variant_weights = {
|
||||||
|
"v1": (0.0, 0.0, 0.0),
|
||||||
|
"v2_a": (5.0, 0.0, 0.0),
|
||||||
|
"v2_b": (5.0, 0.5, 0.0),
|
||||||
|
"v2_c": (5.0, 0.5, 10.0),
|
||||||
|
}
|
||||||
|
if loss_variant not in variant_weights:
|
||||||
|
raise ValueError(f"unknown loss variant: {loss_variant}")
|
||||||
|
lambda_span, lambda_div, lambda_band = variant_weights[loss_variant]
|
||||||
|
grid_size = output.weights["text"].shape[1]
|
||||||
|
if method == "M3":
|
||||||
|
text_reference = torch.bmm(
|
||||||
|
output.weights["text"], sequences["text"].times.unsqueeze(-1)
|
||||||
|
).squeeze(-1)
|
||||||
|
text_reference = text_reference / durations[:, None].clamp_min(1e-8)
|
||||||
|
band_targets = {"audio": text_reference, "vision": text_reference}
|
||||||
|
diversity_modalities = ("audio", "vision")
|
||||||
|
else:
|
||||||
|
centers = (
|
||||||
|
torch.arange(grid_size, dtype=durations.dtype, device=durations.device) + 0.5
|
||||||
|
) / grid_size
|
||||||
|
reference = centers.unsqueeze(0).expand(durations.shape[0], -1)
|
||||||
|
band_targets = {name: reference for name in MODALITIES}
|
||||||
|
diversity_modalities = MODALITIES
|
||||||
|
from .losses import alignment_training_loss
|
||||||
|
|
||||||
|
return alignment_training_loss(
|
||||||
|
output,
|
||||||
|
times,
|
||||||
|
durations,
|
||||||
|
reconstruction,
|
||||||
|
lambda_rec=1.0,
|
||||||
|
lambda_con=1.0,
|
||||||
|
lambda_mono=0.1,
|
||||||
|
lambda_span=lambda_span,
|
||||||
|
lambda_div=lambda_div,
|
||||||
|
lambda_band=lambda_band,
|
||||||
|
epsilon=0.02,
|
||||||
|
minimum_span=0.7,
|
||||||
|
coverage_modalities=diversity_modalities,
|
||||||
|
diversity_modalities=diversity_modalities,
|
||||||
|
diversity_min_separation=6,
|
||||||
|
band_targets=band_targets,
|
||||||
|
band_margin=0.1,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _make_learned_model(
|
||||||
|
method: str,
|
||||||
|
dimensions: Mapping[str, int],
|
||||||
|
grid_size: int,
|
||||||
|
hidden_size: int,
|
||||||
|
heads: int,
|
||||||
|
dropout: float,
|
||||||
|
) -> nn.Module:
|
||||||
|
if method == "M3":
|
||||||
|
return TextAnchoredCrossAttention(
|
||||||
|
dimensions,
|
||||||
|
grid_size=grid_size,
|
||||||
|
hidden_size=hidden_size,
|
||||||
|
heads=heads,
|
||||||
|
dropout=dropout,
|
||||||
|
)
|
||||||
|
if method == "M4":
|
||||||
|
return SharedLatentTimeline(
|
||||||
|
dimensions,
|
||||||
|
grid_size=grid_size,
|
||||||
|
hidden_size=hidden_size,
|
||||||
|
heads=heads,
|
||||||
|
dropout=dropout,
|
||||||
|
)
|
||||||
|
raise ValueError(f"unknown learned method: {method}")
|
||||||
|
|
||||||
|
|
||||||
|
def _fit_learned_model(
|
||||||
|
method: str,
|
||||||
|
train_samples: Sequence[FeatureSample],
|
||||||
|
val_samples: Sequence[FeatureSample],
|
||||||
|
stats: FeatureStats,
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
seed: int,
|
||||||
|
grid_size: int,
|
||||||
|
hidden_size: int,
|
||||||
|
heads: int,
|
||||||
|
dropout: float,
|
||||||
|
batch_size: int,
|
||||||
|
max_epochs: int,
|
||||||
|
patience: int,
|
||||||
|
learning_rate: float,
|
||||||
|
checkpoint_path: Path,
|
||||||
|
loss_variant: str = "v1",
|
||||||
|
) -> tuple[nn.Module, dict[str, Any]]:
|
||||||
|
_seed_everything(seed)
|
||||||
|
dimensions = {name: train_samples[0].features[name].shape[1] for name in MODALITIES}
|
||||||
|
model = _make_learned_model(method, dimensions, grid_size, hidden_size, heads, dropout).to(device)
|
||||||
|
decoder = AlignmentReconstructor(hidden_size, dropout).to(device)
|
||||||
|
optimizer = torch.optim.AdamW(
|
||||||
|
[*model.parameters(), *decoder.parameters()], lr=learning_rate, weight_decay=1e-4
|
||||||
|
)
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
train_mask_generator = torch.Generator(device=device)
|
||||||
|
train_mask_generator.manual_seed(seed + 31)
|
||||||
|
history: list[dict[str, float]] = []
|
||||||
|
best_loss = math.inf
|
||||||
|
best_epoch = 0
|
||||||
|
best_model: dict[str, Tensor] | None = None
|
||||||
|
best_decoder: dict[str, Tensor] | None = None
|
||||||
|
patience_used = 0
|
||||||
|
|
||||||
|
for epoch in range(1, max_epochs + 1):
|
||||||
|
model.train()
|
||||||
|
decoder.train()
|
||||||
|
train_total = 0.0
|
||||||
|
train_count = 0
|
||||||
|
for batch_samples in _batches(train_samples, batch_size, shuffle=True, rng=rng):
|
||||||
|
sequences, durations, _ = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
output = model(sequences)
|
||||||
|
total, _ = _training_objective(
|
||||||
|
output,
|
||||||
|
sequences,
|
||||||
|
durations,
|
||||||
|
decoder,
|
||||||
|
train_mask_generator,
|
||||||
|
method=method,
|
||||||
|
loss_variant=loss_variant,
|
||||||
|
)
|
||||||
|
if not torch.isfinite(total):
|
||||||
|
raise FloatingPointError(f"non-finite {method} objective at epoch {epoch}")
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
total.backward()
|
||||||
|
nn.utils.clip_grad_norm_([*model.parameters(), *decoder.parameters()], 1.0)
|
||||||
|
optimizer.step()
|
||||||
|
train_total += float(total.detach().item()) * len(batch_samples)
|
||||||
|
train_count += len(batch_samples)
|
||||||
|
|
||||||
|
model.eval()
|
||||||
|
decoder.eval()
|
||||||
|
val_generator = torch.Generator(device=device)
|
||||||
|
val_generator.manual_seed(seed + 99991)
|
||||||
|
val_total = 0.0
|
||||||
|
val_count = 0
|
||||||
|
val_metric_sums: dict[str, float] = defaultdict(float)
|
||||||
|
with torch.no_grad():
|
||||||
|
for batch_samples in _batches(val_samples, batch_size, shuffle=False, rng=rng):
|
||||||
|
sequences, durations, _ = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
output = model(sequences)
|
||||||
|
total, parts = _training_objective(
|
||||||
|
output,
|
||||||
|
sequences,
|
||||||
|
durations,
|
||||||
|
decoder,
|
||||||
|
val_generator,
|
||||||
|
method=method,
|
||||||
|
loss_variant=loss_variant,
|
||||||
|
)
|
||||||
|
count = len(batch_samples)
|
||||||
|
val_total += float(total.item()) * count
|
||||||
|
val_count += count
|
||||||
|
for key, value in parts.items():
|
||||||
|
val_metric_sums[key] += float(value.item()) * count
|
||||||
|
for name in MODALITIES:
|
||||||
|
val_metric_sums[f"c_row_{name}"] += float(
|
||||||
|
attention_row_similarity(output.weights[name]).mean().item()
|
||||||
|
) * count
|
||||||
|
val_metric_sums[f"c_far_{name}"] += float(
|
||||||
|
attention_row_similarity(output.weights[name], min_separation=6)
|
||||||
|
.mean()
|
||||||
|
.item()
|
||||||
|
) * count
|
||||||
|
val_mean = val_total / max(val_count, 1)
|
||||||
|
history.append(
|
||||||
|
{
|
||||||
|
"epoch": float(epoch),
|
||||||
|
"train_total": train_total / max(train_count, 1),
|
||||||
|
"validation_total": val_mean,
|
||||||
|
**{
|
||||||
|
f"validation_{key}": value / max(val_count, 1)
|
||||||
|
for key, value in val_metric_sums.items()
|
||||||
|
},
|
||||||
|
}
|
||||||
|
)
|
||||||
|
if val_mean < best_loss - 1e-6:
|
||||||
|
best_loss = val_mean
|
||||||
|
best_epoch = epoch
|
||||||
|
best_model = {key: value.detach().cpu().clone() for key, value in model.state_dict().items()}
|
||||||
|
best_decoder = {
|
||||||
|
key: value.detach().cpu().clone() for key, value in decoder.state_dict().items()
|
||||||
|
}
|
||||||
|
patience_used = 0
|
||||||
|
else:
|
||||||
|
patience_used += 1
|
||||||
|
if patience_used >= patience:
|
||||||
|
break
|
||||||
|
|
||||||
|
if best_model is None or best_decoder is None:
|
||||||
|
raise RuntimeError(f"{method} training did not produce a finite checkpoint")
|
||||||
|
model.load_state_dict(best_model)
|
||||||
|
decoder.load_state_dict(best_decoder)
|
||||||
|
model.eval()
|
||||||
|
decoder.eval()
|
||||||
|
checkpoint_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
torch.save(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"seed": seed,
|
||||||
|
"loss_variant": loss_variant,
|
||||||
|
"best_epoch": best_epoch,
|
||||||
|
"best_validation_objective": best_loss,
|
||||||
|
"model_state_dict": best_model,
|
||||||
|
"training_decoder_state_dict": best_decoder,
|
||||||
|
"history": history,
|
||||||
|
},
|
||||||
|
checkpoint_path,
|
||||||
|
)
|
||||||
|
return model, {
|
||||||
|
"best_epoch": best_epoch,
|
||||||
|
"best_validation_objective": best_loss,
|
||||||
|
"history": history,
|
||||||
|
"best_validation_metrics": history[best_epoch - 1],
|
||||||
|
"checkpoint": str(checkpoint_path),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _baseline_output(
|
||||||
|
method: str,
|
||||||
|
sequences: Mapping[str, Any],
|
||||||
|
durations: Tensor,
|
||||||
|
word_intervals: Sequence[Tensor],
|
||||||
|
grid_size: int,
|
||||||
|
) -> AlignmentOutput:
|
||||||
|
if method == "M1":
|
||||||
|
return align_forced_timestamps(sequences, word_intervals, grid_size)
|
||||||
|
if method == "M2":
|
||||||
|
return align_fixed_windows(sequences, durations, grid_size)
|
||||||
|
raise ValueError(f"not a fixed baseline: {method}")
|
||||||
|
|
||||||
|
|
||||||
|
def _alignment_rows(
|
||||||
|
method: str,
|
||||||
|
seed_label: str,
|
||||||
|
fold: int,
|
||||||
|
samples: Sequence[FeatureSample],
|
||||||
|
output: AlignmentOutput,
|
||||||
|
device: torch.device,
|
||||||
|
epsilon: float,
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
rows = []
|
||||||
|
for index, sample in enumerate(samples):
|
||||||
|
for name in MODALITIES:
|
||||||
|
length = len(sample.times[name])
|
||||||
|
weights = output.weights[name][index : index + 1, :, :length]
|
||||||
|
times = torch.as_tensor(sample.times[name], dtype=torch.float32, device=device)[None]
|
||||||
|
valid = torch.as_tensor(sample.valid[name], dtype=torch.bool, device=device)[None]
|
||||||
|
duration = torch.tensor([sample.duration_s], dtype=torch.float32, device=device)
|
||||||
|
trajectory = alignment_trajectory(weights, times, duration)
|
||||||
|
mvr = monotonicity_violation_rate(trajectory, epsilon)
|
||||||
|
entropy = normalized_attention_entropy(weights, valid)
|
||||||
|
width = attention_width80(weights)
|
||||||
|
rows.append(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"seed": seed_label,
|
||||||
|
"fold": fold,
|
||||||
|
"sample_id": sample.sample_id,
|
||||||
|
"modality": name,
|
||||||
|
"mvr": float(mvr[0].item()),
|
||||||
|
"normalized_entropy": float(entropy.mean().item()),
|
||||||
|
"width80_source_positions": float(width.float().mean().item()),
|
||||||
|
"c_row": float(attention_row_similarity(weights).mean().item()),
|
||||||
|
"c_far": float(
|
||||||
|
attention_row_similarity(weights, min_separation=6).mean().item()
|
||||||
|
),
|
||||||
|
"expected_time_start_s": float(trajectory[0, 0].item() * sample.duration_s),
|
||||||
|
"expected_time_end_s": float(trajectory[0, -1].item() * sample.duration_s),
|
||||||
|
"trajectory_span_fraction": float((trajectory[0, -1] - trajectory[0, 0]).item()),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def _collect_representations(
|
||||||
|
method: str,
|
||||||
|
samples: Sequence[FeatureSample],
|
||||||
|
stats: FeatureStats,
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
grid_size: int,
|
||||||
|
batch_size: int,
|
||||||
|
model: nn.Module | None = None,
|
||||||
|
) -> tuple[dict[str, dict[str, np.ndarray]], dict[str, dict[str, np.ndarray]]]:
|
||||||
|
aligned_by_id: dict[str, dict[str, np.ndarray]] = {}
|
||||||
|
weights_by_id: dict[str, dict[str, np.ndarray]] = {}
|
||||||
|
rng = np.random.default_rng(0)
|
||||||
|
if model is not None:
|
||||||
|
model.eval()
|
||||||
|
with torch.no_grad():
|
||||||
|
for batch_samples in _batches(samples, batch_size, shuffle=False, rng=rng):
|
||||||
|
sequences, durations, intervals = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
if method in {"M1", "M2"}:
|
||||||
|
output = _baseline_output(method, sequences, durations, intervals, grid_size)
|
||||||
|
elif model is not None:
|
||||||
|
output = model(sequences)
|
||||||
|
else:
|
||||||
|
raise ValueError(f"a trained model is required for {method}")
|
||||||
|
for index, sample in enumerate(batch_samples):
|
||||||
|
aligned: dict[str, np.ndarray] = {}
|
||||||
|
weights: dict[str, np.ndarray] = {}
|
||||||
|
for name in MODALITIES:
|
||||||
|
length = len(sample.features[name])
|
||||||
|
matrix = output.weights[name][index, :, :length]
|
||||||
|
source = sequences[name].features[index, :length]
|
||||||
|
pooled = matrix.to(source.dtype) @ source
|
||||||
|
aligned[name] = pooled.detach().cpu().numpy().astype(np.float32, copy=False)
|
||||||
|
weights[name] = matrix.detach().cpu().numpy().astype(np.float32, copy=False)
|
||||||
|
aligned_by_id[sample.sample_id] = aligned
|
||||||
|
weights_by_id[sample.sample_id] = weights
|
||||||
|
return aligned_by_id, weights_by_id
|
||||||
|
|
||||||
|
|
||||||
|
def _save_alignment(
|
||||||
|
output_dir: Path,
|
||||||
|
method: str,
|
||||||
|
seed_label: str,
|
||||||
|
sample: FeatureSample,
|
||||||
|
weights: Mapping[str, np.ndarray],
|
||||||
|
) -> None:
|
||||||
|
path = output_dir / "alignments" / method / f"seed_{seed_label}" / (
|
||||||
|
sample.sample_id.replace("/", "__") + ".npz"
|
||||||
|
)
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
values: dict[str, np.ndarray] = {"sample_id": np.asarray(sample.sample_id)}
|
||||||
|
for name in MODALITIES:
|
||||||
|
trajectory = weights[name] @ sample.times[name] / max(sample.duration_s, 1e-8)
|
||||||
|
values[f"weights_{name}"] = weights[name]
|
||||||
|
values[f"times_{name}_s"] = sample.times[name].astype(np.float32, copy=False)
|
||||||
|
values[f"valid_{name}"] = sample.valid[name]
|
||||||
|
values[f"trajectory_{name}"] = trajectory.astype(np.float32, copy=False)
|
||||||
|
np.savez_compressed(path, **values)
|
||||||
|
|
||||||
|
|
||||||
|
def _write_csv(path: Path, rows: Sequence[Mapping[str, Any]]) -> None:
|
||||||
|
if not rows:
|
||||||
|
return
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
columns = list(dict.fromkeys(key for row in rows for key in row))
|
||||||
|
with path.open("w", encoding="utf-8-sig", newline="") as file:
|
||||||
|
writer = csv.DictWriter(file, fieldnames=columns, extrasaction="ignore")
|
||||||
|
writer.writeheader()
|
||||||
|
writer.writerows(rows)
|
||||||
|
|
||||||
|
|
||||||
|
def _group_summary(
|
||||||
|
rows: Sequence[Mapping[str, Any]], group_columns: Sequence[str], metric_columns: Sequence[str]
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
groups: dict[tuple[Any, ...], list[Mapping[str, Any]]] = defaultdict(list)
|
||||||
|
for row in rows:
|
||||||
|
groups[tuple(row[column] for column in group_columns)].append(row)
|
||||||
|
output = []
|
||||||
|
for key, values in groups.items():
|
||||||
|
summary: dict[str, Any] = dict(zip(group_columns, key))
|
||||||
|
summary["n"] = len(values)
|
||||||
|
for metric in metric_columns:
|
||||||
|
numbers = [float(row[metric]) for row in values if row.get(metric) not in (None, "")]
|
||||||
|
numbers = [value for value in numbers if math.isfinite(value)]
|
||||||
|
if numbers:
|
||||||
|
summary[f"{metric}_mean"] = statistics.fmean(numbers)
|
||||||
|
summary[f"{metric}_std"] = statistics.stdev(numbers) if len(numbers) > 1 else 0.0
|
||||||
|
output.append(summary)
|
||||||
|
return output
|
||||||
|
|
||||||
|
|
||||||
|
def _comparison_table(summaries: Mapping[str, Sequence[Mapping[str, Any]]]) -> list[dict[str, Any]]:
|
||||||
|
"""Make one compact, multi-metric table without inventing a composite score."""
|
||||||
|
alignment = {(row["method"], row["modality"]): row for row in summaries["alignment"]}
|
||||||
|
retrieval = {(row["method"], row["direction"]): row for row in summaries["retrieval"]}
|
||||||
|
reconstruction = {
|
||||||
|
(row["method"], row["target_modality"]): row for row in summaries["reconstruction"]
|
||||||
|
}
|
||||||
|
emotion = {row["method"]: row for row in summaries["emotion"]}
|
||||||
|
table: list[dict[str, Any]] = []
|
||||||
|
for method in ("M1", "M2", "M3", "M4"):
|
||||||
|
row: dict[str, Any] = {"method": method}
|
||||||
|
for modality in MODALITIES:
|
||||||
|
metrics = alignment[(method, modality)]
|
||||||
|
row[f"mvr_{modality}"] = metrics.get("mvr_mean")
|
||||||
|
row[f"entropy_{modality}"] = metrics.get("normalized_entropy_mean")
|
||||||
|
row[f"trajectory_span_{modality}"] = metrics.get("trajectory_span_fraction_mean")
|
||||||
|
for direction in ("text_to_audio", "text_to_vision"):
|
||||||
|
metrics = retrieval[(method, direction)]
|
||||||
|
row[f"r_at_1_{direction}"] = metrics.get("r_at_1_mean")
|
||||||
|
row[f"r_at_5_{direction}"] = metrics.get("r_at_5_mean")
|
||||||
|
for modality in MODALITIES:
|
||||||
|
metrics = reconstruction[(method, modality)]
|
||||||
|
row[f"reconstruction_mae_{modality}"] = metrics.get("mae_standardized_mean")
|
||||||
|
for metric in ("accuracy", "macro_f1", "mae", "pearson"):
|
||||||
|
row[f"emotion_{metric}"] = emotion[method].get(f"{metric}_mean")
|
||||||
|
table.append(row)
|
||||||
|
return table
|
||||||
|
|
||||||
|
|
||||||
|
def _make_figures(
|
||||||
|
output_dir: Path,
|
||||||
|
example_id: str,
|
||||||
|
example_sample: FeatureSample,
|
||||||
|
example_weights: Mapping[str, Mapping[str, np.ndarray]],
|
||||||
|
grid_size: int,
|
||||||
|
) -> None:
|
||||||
|
method_order = ("M1", "M2", "M3", "M4")
|
||||||
|
fig, axes = plt.subplots(4, 3, figsize=(15, 13), constrained_layout=True)
|
||||||
|
for row, method in enumerate(method_order):
|
||||||
|
if method not in example_weights:
|
||||||
|
continue
|
||||||
|
for col, name in enumerate(MODALITIES):
|
||||||
|
ax = axes[row, col]
|
||||||
|
matrix = example_weights[method][name]
|
||||||
|
image = ax.imshow(matrix, origin="lower", aspect="auto", interpolation="nearest", cmap="magma")
|
||||||
|
ax.set_title(f"{method} · {name}")
|
||||||
|
ax.set_xlabel("source position")
|
||||||
|
ax.set_ylabel("shared grid slot")
|
||||||
|
ax.set_yticks(np.linspace(0, grid_size - 1, 5, dtype=int))
|
||||||
|
fig.colorbar(image, ax=ax, fraction=0.046, pad=0.04)
|
||||||
|
fig.suptitle(f"Alignment matrices on held-out sample {example_id}")
|
||||||
|
fig.savefig(output_dir / "typical_alignment_heatmaps.png", dpi=170)
|
||||||
|
plt.close(fig)
|
||||||
|
|
||||||
|
fig, axes = plt.subplots(2, 2, figsize=(13, 9), constrained_layout=True)
|
||||||
|
x = (np.arange(grid_size, dtype=np.float32) + 0.5) / grid_size
|
||||||
|
for ax, method in zip(axes.flat, method_order):
|
||||||
|
for name in MODALITIES:
|
||||||
|
matrix = example_weights[method][name]
|
||||||
|
duration = example_sample.duration_s
|
||||||
|
trajectory = matrix @ example_sample.times[name] / max(duration, 1e-8)
|
||||||
|
ax.plot(x, trajectory, label=name)
|
||||||
|
ax.plot([0, 1], [0, 1], linestyle="--", color="black", alpha=0.5, label="uniform-time reference")
|
||||||
|
ax.set_title(method)
|
||||||
|
ax.set_xlabel("shared-grid position")
|
||||||
|
ax.set_ylabel("expected source time / clip duration")
|
||||||
|
ax.set_xlim(0, 1)
|
||||||
|
ax.set_ylim(0, 1)
|
||||||
|
ax.grid(alpha=0.2)
|
||||||
|
axes[0, 0].legend(fontsize=8)
|
||||||
|
fig.suptitle(f"Alignment trajectories on held-out sample {example_id}")
|
||||||
|
fig.savefig(output_dir / "typical_alignment_trajectories.png", dpi=170)
|
||||||
|
plt.close(fig)
|
||||||
|
|
||||||
|
|
||||||
|
def run(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
start_time = time.time()
|
||||||
|
_seed_everything(args.seeds[0])
|
||||||
|
if args.device == "auto":
|
||||||
|
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||||
|
else:
|
||||||
|
device = torch.device(args.device)
|
||||||
|
if device.type == "cuda" and not torch.cuda.is_available():
|
||||||
|
raise RuntimeError("CUDA was requested but is not available in this WSL environment")
|
||||||
|
|
||||||
|
samples = load_feature_samples(args.feature_dir, args.manifest)
|
||||||
|
if len(samples) != args.expected_samples:
|
||||||
|
raise ValueError(f"expected {args.expected_samples} samples, found {len(samples)}")
|
||||||
|
groups = [sample.group_id for sample in samples]
|
||||||
|
if len(set(groups)) < args.folds:
|
||||||
|
raise ValueError("fewer video_id groups than requested folds")
|
||||||
|
split_iter = GroupKFold(n_splits=args.folds).split(np.zeros(len(samples)), groups=groups)
|
||||||
|
splits = [(train.tolist(), val.tolist()) for train, val in split_iter]
|
||||||
|
splits_json = [
|
||||||
|
{
|
||||||
|
"fold": fold + 1,
|
||||||
|
"train_sample_ids": [samples[index].sample_id for index in train],
|
||||||
|
"validation_sample_ids": [samples[index].sample_id for index in val],
|
||||||
|
"train_video_ids": sorted({samples[index].group_id for index in train}),
|
||||||
|
"validation_video_ids": sorted({samples[index].group_id for index in val}),
|
||||||
|
}
|
||||||
|
for fold, (train, val) in enumerate(splits)
|
||||||
|
]
|
||||||
|
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
print(
|
||||||
|
f"[start] samples={len(samples)} groups={len(set(groups))} folds={args.folds} "
|
||||||
|
f"seeds={args.seeds} device={device}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
(args.output_dir / "splits.json").write_text(
|
||||||
|
json.dumps(splits_json, ensure_ascii=False, indent=2), encoding="utf-8"
|
||||||
|
)
|
||||||
|
|
||||||
|
alignment_rows: list[dict[str, Any]] = []
|
||||||
|
retrieval_rows: list[dict[str, Any]] = []
|
||||||
|
reconstruction_rows: list[dict[str, Any]] = []
|
||||||
|
emotion_rows: list[dict[str, Any]] = []
|
||||||
|
training_rows: list[dict[str, Any]] = []
|
||||||
|
examples: dict[str, dict[str, Mapping[str, np.ndarray]]] = defaultdict(dict)
|
||||||
|
sample_by_id = {sample.sample_id: sample for sample in samples}
|
||||||
|
preferred_example = args.example_id if args.example_id in sample_by_id else samples[0].sample_id
|
||||||
|
preferred_seed = args.seeds[0]
|
||||||
|
example_sample = sample_by_id[preferred_example]
|
||||||
|
|
||||||
|
for fold_index, (train_indices, val_indices) in enumerate(splits, start=1):
|
||||||
|
train_samples = [samples[index] for index in train_indices]
|
||||||
|
val_samples = [samples[index] for index in val_indices]
|
||||||
|
stats = fit_feature_stats(train_samples)
|
||||||
|
fold_output = args.output_dir / f"fold_{fold_index:02d}"
|
||||||
|
print(
|
||||||
|
f"[fold {fold_index}/{args.folds}] train={len(train_samples)} validation={len(val_samples)} "
|
||||||
|
f"train_video_ids={len({sample.group_id for sample in train_samples})} "
|
||||||
|
f"validation_video_ids={len({sample.group_id for sample in val_samples})}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
# M1 and M2 are deterministic methods with no learned alignment loss.
|
||||||
|
for method in ("M1", "M2"):
|
||||||
|
print(f"[fold {fold_index}] evaluate {method} and train identical probes", flush=True)
|
||||||
|
train_aligned, _ = _collect_representations(
|
||||||
|
method,
|
||||||
|
train_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
val_aligned, val_weights = _collect_representations(
|
||||||
|
method,
|
||||||
|
val_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
combined = {**train_aligned, **val_aligned}
|
||||||
|
rng = np.random.default_rng(args.seeds[0] + fold_index)
|
||||||
|
for batch_samples in _batches(val_samples, args.batch_size, shuffle=False, rng=rng):
|
||||||
|
sequences, durations, intervals = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
output = _baseline_output(method, sequences, durations, intervals, args.grid_size)
|
||||||
|
alignment_rows.extend(
|
||||||
|
_alignment_rows(method, "fixed", fold_index, batch_samples, output, device, args.mvr_epsilon)
|
||||||
|
)
|
||||||
|
for sample in val_samples:
|
||||||
|
_save_alignment(args.output_dir, method, "fixed", sample, val_weights[sample.sample_id])
|
||||||
|
if sample.sample_id == preferred_example:
|
||||||
|
examples[sample.sample_id][method] = val_weights[sample.sample_id]
|
||||||
|
probe_seed = args.seeds[0] + fold_index * 100
|
||||||
|
retrieval_rows.extend(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"seed": "fixed",
|
||||||
|
"fold": fold_index,
|
||||||
|
**row,
|
||||||
|
}
|
||||||
|
for row in run_retrieval_probe(
|
||||||
|
[sample.sample_id for sample in train_samples],
|
||||||
|
[sample.sample_id for sample in val_samples],
|
||||||
|
combined,
|
||||||
|
device=device,
|
||||||
|
seed=probe_seed,
|
||||||
|
epochs=args.retrieval_probe_epochs,
|
||||||
|
batch_size=args.probe_batch_size,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
reconstruction_rows.extend(
|
||||||
|
{"method": method, "seed": "fixed", "fold": fold_index, **row}
|
||||||
|
for row in run_reconstruction_probe(
|
||||||
|
[sample.sample_id for sample in train_samples],
|
||||||
|
[sample.sample_id for sample in val_samples],
|
||||||
|
combined,
|
||||||
|
device=device,
|
||||||
|
seed=probe_seed + 1,
|
||||||
|
ratio=args.mask_ratio,
|
||||||
|
epochs=args.reconstruction_probe_epochs,
|
||||||
|
batch_size=args.probe_batch_size,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
emotion_rows.append(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"seed": "fixed",
|
||||||
|
"fold": fold_index,
|
||||||
|
**run_frozen_emotion_probe(train_samples, val_samples, combined),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
# M3/M4 share one unsupervised objective, split, and training budget.
|
||||||
|
for seed in args.seeds:
|
||||||
|
for method in ("M3", "M4"):
|
||||||
|
print(f"[fold {fold_index}] train {method}, seed={seed}", flush=True)
|
||||||
|
checkpoint = fold_output / f"seed_{seed}" / f"{method}.pt"
|
||||||
|
model, training_info = _fit_learned_model(
|
||||||
|
method,
|
||||||
|
train_samples,
|
||||||
|
val_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
seed=seed + fold_index * 1009,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
hidden_size=args.hidden_size,
|
||||||
|
heads=args.heads,
|
||||||
|
dropout=args.dropout,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
max_epochs=args.epochs,
|
||||||
|
patience=args.patience,
|
||||||
|
learning_rate=args.learning_rate,
|
||||||
|
checkpoint_path=checkpoint,
|
||||||
|
)
|
||||||
|
training_rows.append(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"seed": seed,
|
||||||
|
"fold": fold_index,
|
||||||
|
"best_epoch": training_info["best_epoch"],
|
||||||
|
"best_validation_objective": training_info["best_validation_objective"],
|
||||||
|
"checkpoint": training_info["checkpoint"],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f"[fold {fold_index}] {method}, seed={seed} best_epoch={training_info['best_epoch']} "
|
||||||
|
f"val_objective={training_info['best_validation_objective']:.5f}; running frozen probes",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
train_aligned, _ = _collect_representations(
|
||||||
|
method,
|
||||||
|
train_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
model=model,
|
||||||
|
)
|
||||||
|
val_aligned, val_weights = _collect_representations(
|
||||||
|
method,
|
||||||
|
val_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
model=model,
|
||||||
|
)
|
||||||
|
combined = {**train_aligned, **val_aligned}
|
||||||
|
rng = np.random.default_rng(seed + fold_index)
|
||||||
|
for batch_samples in _batches(val_samples, args.batch_size, shuffle=False, rng=rng):
|
||||||
|
sequences, durations, _ = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
output = model(sequences)
|
||||||
|
alignment_rows.extend(
|
||||||
|
_alignment_rows(method, str(seed), fold_index, batch_samples, output, device, args.mvr_epsilon)
|
||||||
|
)
|
||||||
|
for sample in val_samples:
|
||||||
|
_save_alignment(args.output_dir, method, str(seed), sample, val_weights[sample.sample_id])
|
||||||
|
if sample.sample_id == preferred_example and seed == preferred_seed:
|
||||||
|
examples[sample.sample_id][method] = val_weights[sample.sample_id]
|
||||||
|
|
||||||
|
probe_seed = seed + fold_index * 100 + (3 if method == "M3" else 7)
|
||||||
|
retrieval_rows.extend(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"seed": seed,
|
||||||
|
"fold": fold_index,
|
||||||
|
**row,
|
||||||
|
}
|
||||||
|
for row in run_retrieval_probe(
|
||||||
|
[sample.sample_id for sample in train_samples],
|
||||||
|
[sample.sample_id for sample in val_samples],
|
||||||
|
combined,
|
||||||
|
device=device,
|
||||||
|
seed=probe_seed,
|
||||||
|
epochs=args.retrieval_probe_epochs,
|
||||||
|
batch_size=args.probe_batch_size,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
reconstruction_rows.extend(
|
||||||
|
{"method": method, "seed": seed, "fold": fold_index, **row}
|
||||||
|
for row in run_reconstruction_probe(
|
||||||
|
[sample.sample_id for sample in train_samples],
|
||||||
|
[sample.sample_id for sample in val_samples],
|
||||||
|
combined,
|
||||||
|
device=device,
|
||||||
|
seed=probe_seed + 1,
|
||||||
|
ratio=args.mask_ratio,
|
||||||
|
epochs=args.reconstruction_probe_epochs,
|
||||||
|
batch_size=args.probe_batch_size,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
emotion_rows.append(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"seed": seed,
|
||||||
|
"fold": fold_index,
|
||||||
|
**run_frozen_emotion_probe(train_samples, val_samples, combined),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
del model
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
|
||||||
|
_write_csv(args.output_dir / "alignment_metrics.csv", alignment_rows)
|
||||||
|
_write_csv(args.output_dir / "retrieval_probe_metrics.csv", retrieval_rows)
|
||||||
|
_write_csv(args.output_dir / "reconstruction_probe_metrics.csv", reconstruction_rows)
|
||||||
|
_write_csv(args.output_dir / "frozen_emotion_probe_metrics.csv", emotion_rows)
|
||||||
|
_write_csv(args.output_dir / "training_summary.csv", training_rows)
|
||||||
|
|
||||||
|
summaries = {
|
||||||
|
"alignment": _group_summary(
|
||||||
|
alignment_rows,
|
||||||
|
("method", "modality"),
|
||||||
|
("mvr", "normalized_entropy", "width80_source_positions", "expected_time_start_s", "expected_time_end_s", "trajectory_span_fraction"),
|
||||||
|
),
|
||||||
|
"retrieval": _group_summary(
|
||||||
|
retrieval_rows,
|
||||||
|
("method", "direction"),
|
||||||
|
("r_at_1", "r_at_5", "mrr"),
|
||||||
|
),
|
||||||
|
"reconstruction": _group_summary(
|
||||||
|
reconstruction_rows,
|
||||||
|
("method", "target_modality"),
|
||||||
|
("mae_standardized", "smooth_l1_standardized"),
|
||||||
|
),
|
||||||
|
"emotion": _group_summary(
|
||||||
|
emotion_rows,
|
||||||
|
("method",),
|
||||||
|
("accuracy", "macro_f1", "mae", "pearson"),
|
||||||
|
),
|
||||||
|
}
|
||||||
|
(args.output_dir / "summary.json").write_text(
|
||||||
|
json.dumps(summaries, ensure_ascii=False, indent=2, allow_nan=False), encoding="utf-8"
|
||||||
|
)
|
||||||
|
for name, rows in summaries.items():
|
||||||
|
_write_csv(args.output_dir / f"{name}_summary.csv", rows)
|
||||||
|
comparison_table = _comparison_table(summaries)
|
||||||
|
_write_csv(args.output_dir / "comparison_summary.csv", comparison_table)
|
||||||
|
|
||||||
|
if preferred_example in examples and set(examples[preferred_example]) == {"M1", "M2", "M3", "M4"}:
|
||||||
|
_make_figures(args.output_dir, preferred_example, example_sample, examples[preferred_example], args.grid_size)
|
||||||
|
|
||||||
|
manifest = {
|
||||||
|
"created_utc": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"sample_count": len(samples),
|
||||||
|
"group_count": len(set(groups)),
|
||||||
|
"folds": args.folds,
|
||||||
|
"seeds": args.seeds,
|
||||||
|
"device": str(device),
|
||||||
|
"gpu_name": torch.cuda.get_device_name(device) if device.type == "cuda" else None,
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"parameters": {
|
||||||
|
"grid_size": args.grid_size,
|
||||||
|
"hidden_size": args.hidden_size,
|
||||||
|
"heads": args.heads,
|
||||||
|
"dropout": args.dropout,
|
||||||
|
"batch_size": args.batch_size,
|
||||||
|
"epochs_max": args.epochs,
|
||||||
|
"early_stopping_patience": args.patience,
|
||||||
|
"learning_rate": args.learning_rate,
|
||||||
|
"mask_ratio": args.mask_ratio,
|
||||||
|
"retrieval_probe_epochs": args.retrieval_probe_epochs,
|
||||||
|
"reconstruction_probe_epochs": args.reconstruction_probe_epochs,
|
||||||
|
"mvr_epsilon": args.mvr_epsilon,
|
||||||
|
},
|
||||||
|
"objective": "masked reconstruction + cross-modal contrastive + temporal monotonicity; emotion labels unused",
|
||||||
|
"split_rule": "GroupKFold by group_id/video_id",
|
||||||
|
"example_sample_id": preferred_example,
|
||||||
|
"elapsed_seconds": time.time() - start_time,
|
||||||
|
"interpretation_limits": [
|
||||||
|
"Grid-index retrieval is a representation-consistency probe, not independent temporal ground truth.",
|
||||||
|
"Masked reconstruction uses a decoder trained on the training fold and reports standardized-feature errors.",
|
||||||
|
"The emotion probe is a small-sample downstream utility check, not a claim of generalization to MOSEI.",
|
||||||
|
"No human event timestamps are available, so human IoU/MATE is not reported.",
|
||||||
|
],
|
||||||
|
}
|
||||||
|
(args.output_dir / "run_manifest.json").write_text(
|
||||||
|
json.dumps(manifest, ensure_ascii=False, indent=2), encoding="utf-8"
|
||||||
|
)
|
||||||
|
(args.output_dir / "README.md").write_text(
|
||||||
|
"# Q1 method comparison\n\n"
|
||||||
|
"This folder contains grouped cross-validation results for M1–M4. M1/M2 are fixed rules; M3/M4 are trained without emotion labels. "
|
||||||
|
"All learned models and probes use training-fold-only feature normalization, and folds are grouped by `video_id`.\n\n"
|
||||||
|
"`alignment_metrics.csv` reports expected-time trajectories, monotonicity, attention entropy, and width diagnostics. "
|
||||||
|
"`retrieval_probe_metrics.csv` uses a separately trained linear projection probe; its grid-index positives are not independent temporal ground truth. "
|
||||||
|
"`reconstruction_probe_metrics.csv` reports held-out masked reconstruction error in training-fold standardized feature units. "
|
||||||
|
"`frozen_emotion_probe_metrics.csv` is a small-sample downstream utility check.\n\n"
|
||||||
|
"No human event-time annotation is present, so the results cannot establish direct human alignment accuracy. "
|
||||||
|
"See `run_manifest.json` for parameters, seeds, device, and interpretation limits.\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f"[done] elapsed_seconds={manifest['elapsed_seconds']:.1f} output={args.output_dir}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
return manifest
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
project_dir = Path(__file__).resolve().parents[1]
|
||||||
|
parser = argparse.ArgumentParser(description="Compare Q1 M1-M4 alignment methods with grouped CV.")
|
||||||
|
parser.add_argument("--feature-dir", type=Path, default=project_dir / "outputs/q1_features/features")
|
||||||
|
parser.add_argument("--manifest", type=Path, default=project_dir / "outputs/audit/manifest.csv")
|
||||||
|
parser.add_argument("--output-dir", type=Path, default=project_dir / "outputs/method_comparison")
|
||||||
|
parser.add_argument("--device", default="auto", help="auto, cpu, or a torch device such as cuda:0")
|
||||||
|
parser.add_argument("--expected-samples", type=int, default=100)
|
||||||
|
parser.add_argument("--folds", type=int, default=5)
|
||||||
|
parser.add_argument("--seeds", type=int, nargs="+", default=[42, 3407, 2026])
|
||||||
|
parser.add_argument("--grid-size", type=int, default=50)
|
||||||
|
parser.add_argument("--hidden-size", type=int, default=128)
|
||||||
|
parser.add_argument("--heads", type=int, default=4)
|
||||||
|
parser.add_argument("--dropout", type=float, default=0.1)
|
||||||
|
parser.add_argument("--batch-size", type=int, default=8)
|
||||||
|
parser.add_argument("--probe-batch-size", type=int, default=16)
|
||||||
|
parser.add_argument("--epochs", type=int, default=50)
|
||||||
|
parser.add_argument("--patience", type=int, default=8)
|
||||||
|
parser.add_argument("--learning-rate", type=float, default=1e-4)
|
||||||
|
parser.add_argument("--mask-ratio", type=float, default=0.2)
|
||||||
|
parser.add_argument("--retrieval-probe-epochs", type=int, default=20)
|
||||||
|
parser.add_argument("--reconstruction-probe-epochs", type=int, default=25)
|
||||||
|
parser.add_argument("--mvr-epsilon", type=float, default=0.02)
|
||||||
|
parser.add_argument("--example-id", default="-tPCytz4rww/12")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
manifest = run(args)
|
||||||
|
print(json.dumps(manifest, ensure_ascii=False, indent=2))
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,773 @@
|
|||||||
|
"""Evaluate same-time cross-modal correspondence against within-clip shifts.
|
||||||
|
|
||||||
|
The alignment model is frozen. A small train-only linear projection probe maps
|
||||||
|
the aligned raw features into a shared space using diagonal-versus-off-diagonal
|
||||||
|
InfoNCE within each training clip. All reported correspondence scores are then
|
||||||
|
computed on held-out video_id groups.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
import platform
|
||||||
|
import random
|
||||||
|
import time
|
||||||
|
from collections import defaultdict
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Mapping, Sequence
|
||||||
|
|
||||||
|
import matplotlib
|
||||||
|
|
||||||
|
matplotlib.use("Agg")
|
||||||
|
import matplotlib.pyplot as plt
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
import torch.nn.functional as F
|
||||||
|
from sklearn.metrics import roc_auc_score
|
||||||
|
from torch import Tensor, nn
|
||||||
|
|
||||||
|
from .compare_methods import _collect_representations
|
||||||
|
from .experiment_data import (
|
||||||
|
FeatureSample,
|
||||||
|
FeatureStats,
|
||||||
|
collate_feature_samples,
|
||||||
|
fit_feature_stats,
|
||||||
|
load_feature_samples,
|
||||||
|
)
|
||||||
|
from .models import SharedLatentTimeline, TextAnchoredCrossAttention
|
||||||
|
from .types import MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
GRID_SIZE = 50
|
||||||
|
HIDDEN_SIZE = 128
|
||||||
|
HEADS = 4
|
||||||
|
OUTPUT_SIZE = 64
|
||||||
|
PAIRINGS = (("text", "audio"), ("text", "vision"), ("audio", "vision"))
|
||||||
|
METHODS = ("M1", "M2", "M3_noSourceTime", "M3_sourceTime", "M4_noSourceTime", "M4_sourceTime")
|
||||||
|
CURVE_MAX_SHIFT = 10
|
||||||
|
NEGATIVE_RADIUS = 2
|
||||||
|
|
||||||
|
|
||||||
|
class CorrespondenceProjection(nn.Module):
|
||||||
|
"""Equal-output-size modality heads used only as a frozen-feature probe."""
|
||||||
|
|
||||||
|
def __init__(self, dimensions: Mapping[str, int], output_size: int = OUTPUT_SIZE) -> None:
|
||||||
|
super().__init__()
|
||||||
|
self.projections = nn.ModuleDict(
|
||||||
|
{name: nn.Linear(dimensions[name], output_size, bias=False) for name in MODALITIES}
|
||||||
|
)
|
||||||
|
|
||||||
|
def forward(self, values: Mapping[str, Tensor]) -> dict[str, Tensor]:
|
||||||
|
return {
|
||||||
|
name: F.normalize(self.projections[name](values[name]), dim=-1)
|
||||||
|
for name in MODALITIES
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _write_csv(path: Path, rows: Sequence[Mapping[str, Any]]) -> None:
|
||||||
|
if not rows:
|
||||||
|
return
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
fields = list(dict.fromkeys(key for row in rows for key in row))
|
||||||
|
with path.open("w", newline="", encoding="utf-8-sig") as handle:
|
||||||
|
writer = csv.DictWriter(handle, fieldnames=fields, extrasaction="ignore")
|
||||||
|
writer.writeheader()
|
||||||
|
writer.writerows(rows)
|
||||||
|
|
||||||
|
|
||||||
|
def _batches(samples: Sequence[FeatureSample], batch_size: int):
|
||||||
|
for start in range(0, len(samples), batch_size):
|
||||||
|
yield list(samples[start : start + batch_size])
|
||||||
|
|
||||||
|
|
||||||
|
def _make_model(method: str, dimensions: Mapping[str, int], source_time: bool) -> nn.Module:
|
||||||
|
if method == "M3":
|
||||||
|
return TextAnchoredCrossAttention(
|
||||||
|
dimensions,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
hidden_size=HIDDEN_SIZE,
|
||||||
|
heads=HEADS,
|
||||||
|
dropout=0.0,
|
||||||
|
source_time_encoding=source_time,
|
||||||
|
)
|
||||||
|
return SharedLatentTimeline(
|
||||||
|
dimensions,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
hidden_size=HIDDEN_SIZE,
|
||||||
|
heads=HEADS,
|
||||||
|
dropout=0.0,
|
||||||
|
absolute_position_encoding=True,
|
||||||
|
source_time_encoding=source_time,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _checkpoint_path(root: Path, fold: int, variant: str) -> Path:
|
||||||
|
if fold == 1:
|
||||||
|
return root / variant / "checkpoint.pt"
|
||||||
|
return root / f"fold_{fold:02d}" / variant / "checkpoint.pt"
|
||||||
|
|
||||||
|
|
||||||
|
def _collect_learned_representations(
|
||||||
|
*,
|
||||||
|
method_name: str,
|
||||||
|
fold: int,
|
||||||
|
train_samples: Sequence[FeatureSample],
|
||||||
|
validation_samples: Sequence[FeatureSample],
|
||||||
|
stats: FeatureStats,
|
||||||
|
checkpoint_root: Path,
|
||||||
|
device: torch.device,
|
||||||
|
batch_size: int,
|
||||||
|
) -> dict[str, dict[str, np.ndarray]]:
|
||||||
|
method = method_name[:2]
|
||||||
|
source_time = method_name.endswith("_sourceTime")
|
||||||
|
checkpoint_path = _checkpoint_path(checkpoint_root, fold, method_name)
|
||||||
|
if not checkpoint_path.is_file():
|
||||||
|
raise FileNotFoundError(
|
||||||
|
f"missing {method_name} fold {fold} checkpoint: {checkpoint_path}; "
|
||||||
|
"run q1.alignment_heldout_debug for this fold first"
|
||||||
|
)
|
||||||
|
checkpoint = torch.load(checkpoint_path, map_location=device, weights_only=False)
|
||||||
|
if checkpoint.get("variant") != method_name:
|
||||||
|
raise ValueError(f"checkpoint variant does not match {method_name}: {checkpoint_path}")
|
||||||
|
if bool(checkpoint.get("source_time_encoding")) != source_time:
|
||||||
|
raise ValueError(f"checkpoint source-time setting does not match {method_name}: {checkpoint_path}")
|
||||||
|
expected_train = {sample.sample_id for sample in train_samples}
|
||||||
|
expected_validation = {sample.sample_id for sample in validation_samples}
|
||||||
|
if set(checkpoint.get("train_sample_ids", [])) != expected_train:
|
||||||
|
raise ValueError(f"checkpoint training IDs do not match split for {method_name}, fold {fold}")
|
||||||
|
if set(checkpoint.get("validation_sample_ids", [])) != expected_validation:
|
||||||
|
raise ValueError(f"checkpoint validation IDs do not match split for {method_name}, fold {fold}")
|
||||||
|
|
||||||
|
dimensions = {name: train_samples[0].features[name].shape[1] for name in MODALITIES}
|
||||||
|
model = _make_model(method, dimensions, source_time).to(device)
|
||||||
|
model.load_state_dict(checkpoint["model_state_dict"], strict=True)
|
||||||
|
model.eval()
|
||||||
|
result: dict[str, dict[str, np.ndarray]] = {}
|
||||||
|
with torch.no_grad():
|
||||||
|
for batch_samples in _batches([*train_samples, *validation_samples], batch_size):
|
||||||
|
sequences, durations, _ = collate_feature_samples(batch_samples, stats, device)
|
||||||
|
output = model(sequences, durations) if source_time else model(sequences)
|
||||||
|
for index, sample in enumerate(batch_samples):
|
||||||
|
sample_values: dict[str, np.ndarray] = {}
|
||||||
|
for name in MODALITIES:
|
||||||
|
length = len(sample.features[name])
|
||||||
|
weights = output.weights[name][index, :, :length]
|
||||||
|
source = sequences[name].features[index, :length]
|
||||||
|
pooled = weights.to(source.dtype) @ source
|
||||||
|
if pooled.shape[0] != GRID_SIZE:
|
||||||
|
raise ValueError(f"unexpected grid size for {sample.sample_id}/{name}")
|
||||||
|
sample_values[name] = pooled.detach().cpu().numpy().astype(np.float32, copy=False)
|
||||||
|
result[sample.sample_id] = sample_values
|
||||||
|
del model
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
def _stack_ids(
|
||||||
|
sample_ids: Sequence[str],
|
||||||
|
aligned_by_id: Mapping[str, Mapping[str, np.ndarray]],
|
||||||
|
device: torch.device,
|
||||||
|
) -> dict[str, Tensor]:
|
||||||
|
return {
|
||||||
|
name: torch.as_tensor(
|
||||||
|
np.stack([aligned_by_id[sample_id][name] for sample_id in sample_ids]),
|
||||||
|
dtype=torch.float32,
|
||||||
|
device=device,
|
||||||
|
)
|
||||||
|
for name in MODALITIES
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _within_clip_infonce(
|
||||||
|
projected: Mapping[str, Tensor], temperature: float
|
||||||
|
) -> Tensor:
|
||||||
|
"""Symmetric diagonal-vs-shifted loss; negatives stay inside each clip."""
|
||||||
|
batch_size, grid_size = projected["text"].shape[:2]
|
||||||
|
target = torch.arange(grid_size, device=projected["text"].device).repeat(batch_size)
|
||||||
|
losses = []
|
||||||
|
for left_index, left_name in enumerate(MODALITIES):
|
||||||
|
for right_name in MODALITIES[left_index + 1 :]:
|
||||||
|
scores = torch.einsum(
|
||||||
|
"bkd,bld->bkl", projected[left_name], projected[right_name]
|
||||||
|
) / temperature
|
||||||
|
forward = F.cross_entropy(scores.reshape(-1, grid_size), target)
|
||||||
|
reverse = F.cross_entropy(scores.transpose(1, 2).reshape(-1, grid_size), target)
|
||||||
|
losses.append((forward + reverse) / 2)
|
||||||
|
return torch.stack(losses).mean()
|
||||||
|
|
||||||
|
|
||||||
|
def _fit_probe(
|
||||||
|
train_ids: Sequence[str],
|
||||||
|
aligned_by_id: Mapping[str, Mapping[str, np.ndarray]],
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
seed: int,
|
||||||
|
epochs: int,
|
||||||
|
batch_size: int,
|
||||||
|
learning_rate: float,
|
||||||
|
temperature: float,
|
||||||
|
) -> tuple[CorrespondenceProjection, list[dict[str, Any]]]:
|
||||||
|
random.seed(seed)
|
||||||
|
np.random.seed(seed)
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.manual_seed_all(seed)
|
||||||
|
train = _stack_ids(train_ids, aligned_by_id, device)
|
||||||
|
dimensions = {name: int(train[name].shape[-1]) for name in MODALITIES}
|
||||||
|
model = CorrespondenceProjection(dimensions).to(device)
|
||||||
|
optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate, weight_decay=1e-4)
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
history: list[dict[str, Any]] = []
|
||||||
|
model.train()
|
||||||
|
for epoch in range(1, epochs + 1):
|
||||||
|
order = rng.permutation(len(train_ids))
|
||||||
|
loss_sum = 0.0
|
||||||
|
batches = 0
|
||||||
|
for start in range(0, len(order), batch_size):
|
||||||
|
indexes = torch.as_tensor(order[start : start + batch_size], device=device)
|
||||||
|
batch = {name: value.index_select(0, indexes) for name, value in train.items()}
|
||||||
|
loss = _within_clip_infonce(model(batch), temperature)
|
||||||
|
if not torch.isfinite(loss):
|
||||||
|
raise FloatingPointError(f"non-finite correspondence probe loss at epoch {epoch}")
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
loss.backward()
|
||||||
|
nn.utils.clip_grad_norm_(model.parameters(), 1.0)
|
||||||
|
optimizer.step()
|
||||||
|
loss_sum += float(loss.detach().item())
|
||||||
|
batches += 1
|
||||||
|
history.append({"epoch": epoch, "train_loss": loss_sum / max(batches, 1)})
|
||||||
|
return model, history
|
||||||
|
|
||||||
|
|
||||||
|
def _shifted_scores(scores: np.ndarray, delta: int) -> np.ndarray:
|
||||||
|
grid_size = scores.shape[0]
|
||||||
|
if delta >= 0:
|
||||||
|
indexes = np.arange(0, grid_size - delta)
|
||||||
|
return scores[indexes, indexes + delta]
|
||||||
|
indexes = np.arange(-delta, grid_size)
|
||||||
|
return scores[indexes, indexes + delta]
|
||||||
|
|
||||||
|
|
||||||
|
def _sample_metrics(
|
||||||
|
*,
|
||||||
|
method: str,
|
||||||
|
fold: int,
|
||||||
|
sample: FeatureSample,
|
||||||
|
projected: Mapping[str, Tensor],
|
||||||
|
curve_rows: list[dict[str, Any]],
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
rows = []
|
||||||
|
for left_name, right_name in PAIRINGS:
|
||||||
|
left = projected[left_name]
|
||||||
|
right = projected[right_name]
|
||||||
|
score = (left @ right.T).detach().cpu().numpy()
|
||||||
|
grid_size = score.shape[0]
|
||||||
|
if score.shape != (GRID_SIZE, GRID_SIZE):
|
||||||
|
raise ValueError(f"expected {GRID_SIZE}x{GRID_SIZE} score matrix")
|
||||||
|
diagonal = np.diag(score)
|
||||||
|
negative_mask = np.abs(np.arange(grid_size)[:, None] - np.arange(grid_size)[None, :]) > NEGATIVE_RADIUS
|
||||||
|
auc = float(
|
||||||
|
roc_auc_score(
|
||||||
|
np.concatenate((np.ones(grid_size), np.zeros(int(negative_mask.sum())))),
|
||||||
|
np.concatenate((diagonal, score[negative_mask])),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
far_shifts = [
|
||||||
|
_shifted_scores(score, delta).mean()
|
||||||
|
for delta in range(-CURVE_MAX_SHIFT, CURVE_MAX_SHIFT + 1)
|
||||||
|
if abs(delta) > NEGATIVE_RADIUS
|
||||||
|
]
|
||||||
|
row: dict[str, Any] = {
|
||||||
|
"method": method,
|
||||||
|
"fold": fold,
|
||||||
|
"sample_id": sample.sample_id,
|
||||||
|
"video_id": sample.group_id,
|
||||||
|
"pair": f"{left_name}_{right_name}",
|
||||||
|
"same_time_similarity": float(diagonal.mean()),
|
||||||
|
"shifted_far_similarity": float(np.mean(far_shifts)),
|
||||||
|
"same_minus_shifted_margin": float(diagonal.mean() - np.mean(far_shifts)),
|
||||||
|
"matched_vs_shifted_auc": auc,
|
||||||
|
}
|
||||||
|
for direction, directed_score in (
|
||||||
|
(f"{left_name}_to_{right_name}", score),
|
||||||
|
(f"{right_name}_to_{left_name}", score.T),
|
||||||
|
):
|
||||||
|
prediction = directed_score.argmax(axis=1)
|
||||||
|
error = np.abs(prediction - np.arange(grid_size))
|
||||||
|
row[f"exact_r1_{direction}"] = float(np.mean(error == 0))
|
||||||
|
row[f"within_pm1_r1_{direction}"] = float(np.mean(error <= 1))
|
||||||
|
row[f"mase_slots_{direction}"] = float(np.mean(error))
|
||||||
|
|
||||||
|
for delta in range(-CURVE_MAX_SHIFT, CURVE_MAX_SHIFT + 1):
|
||||||
|
curve_rows.append(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"fold": fold,
|
||||||
|
"sample_id": sample.sample_id,
|
||||||
|
"video_id": sample.group_id,
|
||||||
|
"pair": f"{left_name}_{right_name}",
|
||||||
|
"delta": delta,
|
||||||
|
"similarity": float(_shifted_scores(score, delta).mean()),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
rows.append(row)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def _cluster_bootstrap(
|
||||||
|
rows: Sequence[Mapping[str, Any]],
|
||||||
|
metric: str,
|
||||||
|
*,
|
||||||
|
seed: int,
|
||||||
|
repetitions: int = 2000,
|
||||||
|
) -> tuple[float, float, float, int]:
|
||||||
|
grouped: dict[str, list[float]] = defaultdict(list)
|
||||||
|
for row in rows:
|
||||||
|
value = float(row[metric])
|
||||||
|
if np.isfinite(value):
|
||||||
|
grouped[str(row["video_id"])].append(value)
|
||||||
|
groups = np.asarray([np.mean(values) for values in grouped.values()], dtype=np.float64)
|
||||||
|
if groups.size == 0:
|
||||||
|
return float("nan"), float("nan"), float("nan"), 0
|
||||||
|
mean = float(groups.mean())
|
||||||
|
if groups.size == 1:
|
||||||
|
return mean, mean, mean, 1
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
indexes = rng.integers(0, groups.size, size=(repetitions, groups.size))
|
||||||
|
boot = groups[indexes].mean(axis=1)
|
||||||
|
low, high = np.quantile(boot, [0.025, 0.975])
|
||||||
|
return mean, float(low), float(high), int(groups.size)
|
||||||
|
|
||||||
|
|
||||||
|
def _summaries(
|
||||||
|
metric_rows: Sequence[Mapping[str, Any]],
|
||||||
|
curve_rows: Sequence[Mapping[str, Any]],
|
||||||
|
*,
|
||||||
|
seed: int,
|
||||||
|
) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
|
||||||
|
metrics = (
|
||||||
|
"same_time_similarity",
|
||||||
|
"shifted_far_similarity",
|
||||||
|
"same_minus_shifted_margin",
|
||||||
|
"matched_vs_shifted_auc",
|
||||||
|
"exact_r1_text_to_audio",
|
||||||
|
"within_pm1_r1_text_to_audio",
|
||||||
|
"mase_slots_text_to_audio",
|
||||||
|
"exact_r1_audio_to_text",
|
||||||
|
"within_pm1_r1_audio_to_text",
|
||||||
|
"mase_slots_audio_to_text",
|
||||||
|
"exact_r1_text_to_vision",
|
||||||
|
"within_pm1_r1_text_to_vision",
|
||||||
|
"mase_slots_text_to_vision",
|
||||||
|
"exact_r1_vision_to_text",
|
||||||
|
"within_pm1_r1_vision_to_text",
|
||||||
|
"mase_slots_vision_to_text",
|
||||||
|
"exact_r1_audio_to_vision",
|
||||||
|
"within_pm1_r1_audio_to_vision",
|
||||||
|
"mase_slots_audio_to_vision",
|
||||||
|
"exact_r1_vision_to_audio",
|
||||||
|
"within_pm1_r1_vision_to_audio",
|
||||||
|
"mase_slots_vision_to_audio",
|
||||||
|
)
|
||||||
|
grouped: dict[tuple[str, str], list[Mapping[str, Any]]] = defaultdict(list)
|
||||||
|
for row in metric_rows:
|
||||||
|
grouped[(str(row["method"]), str(row["pair"]))].append(row)
|
||||||
|
summary_rows = []
|
||||||
|
for (method, pair), rows in sorted(grouped.items()):
|
||||||
|
result: dict[str, Any] = {
|
||||||
|
"method": method,
|
||||||
|
"pair": pair,
|
||||||
|
"clip_count": len(rows),
|
||||||
|
"video_id_count": len({str(row["video_id"]) for row in rows}),
|
||||||
|
}
|
||||||
|
for metric_index, metric in enumerate(metrics):
|
||||||
|
if metric not in rows[0]:
|
||||||
|
continue
|
||||||
|
mean, low, high, _ = _cluster_bootstrap(
|
||||||
|
rows,
|
||||||
|
metric,
|
||||||
|
seed=seed + metric_index + sum(ord(character) for character in method + pair),
|
||||||
|
)
|
||||||
|
result[f"{metric}_mean"] = mean
|
||||||
|
result[f"{metric}_ci95_low"] = low
|
||||||
|
result[f"{metric}_ci95_high"] = high
|
||||||
|
curve_for_group = [
|
||||||
|
row for row in curve_rows if row["method"] == method and row["pair"] == pair
|
||||||
|
]
|
||||||
|
curve_by_delta: dict[int, list[dict[str, Any]]] = defaultdict(list)
|
||||||
|
for row in curve_for_group:
|
||||||
|
curve_by_delta[int(row["delta"])].append(dict(row))
|
||||||
|
mean_curve = {
|
||||||
|
delta: _cluster_bootstrap(
|
||||||
|
values,
|
||||||
|
"similarity",
|
||||||
|
seed=seed + delta + sum(ord(character) for character in method + pair),
|
||||||
|
)[0]
|
||||||
|
for delta, values in curve_by_delta.items()
|
||||||
|
}
|
||||||
|
if mean_curve:
|
||||||
|
result["peak_delta"] = max(mean_curve, key=mean_curve.get)
|
||||||
|
result["peak_similarity"] = mean_curve[result["peak_delta"]]
|
||||||
|
summary_rows.append(result)
|
||||||
|
|
||||||
|
curve_summary = []
|
||||||
|
curve_groups: dict[tuple[str, str, int], list[Mapping[str, Any]]] = defaultdict(list)
|
||||||
|
for row in curve_rows:
|
||||||
|
curve_groups[(str(row["method"]), str(row["pair"]), int(row["delta"]))].append(row)
|
||||||
|
for (method, pair, delta), rows in sorted(curve_groups.items()):
|
||||||
|
mean, low, high, group_count = _cluster_bootstrap(
|
||||||
|
rows,
|
||||||
|
"similarity",
|
||||||
|
seed=seed + delta + sum(ord(character) for character in method + pair),
|
||||||
|
)
|
||||||
|
curve_summary.append(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"pair": pair,
|
||||||
|
"delta": delta,
|
||||||
|
"mean_similarity": mean,
|
||||||
|
"ci95_low": low,
|
||||||
|
"ci95_high": high,
|
||||||
|
"video_id_count": group_count,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return summary_rows, curve_summary
|
||||||
|
|
||||||
|
|
||||||
|
def _summarize_time_localization(
|
||||||
|
checkpoint_root: Path, fold_count: int, *, seed: int
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
grouped: dict[tuple[str, str], list[dict[str, Any]]] = defaultdict(list)
|
||||||
|
for fold in range(1, fold_count + 1):
|
||||||
|
metrics_path = (
|
||||||
|
checkpoint_root / "per_sample_metrics.csv"
|
||||||
|
if fold == 1
|
||||||
|
else checkpoint_root / f"fold_{fold:02d}" / "per_sample_metrics.csv"
|
||||||
|
)
|
||||||
|
if not metrics_path.is_file():
|
||||||
|
raise FileNotFoundError(f"missing D5 per-sample metrics for fold {fold}: {metrics_path}")
|
||||||
|
with metrics_path.open("r", newline="", encoding="utf-8-sig") as handle:
|
||||||
|
for source in csv.DictReader(handle):
|
||||||
|
sample_id = source["sample_id"]
|
||||||
|
row: dict[str, Any] = {
|
||||||
|
"video_id": sample_id.split("/", 1)[0],
|
||||||
|
"sample_id": sample_id,
|
||||||
|
"variant": source["variant"],
|
||||||
|
"modality": source["modality"],
|
||||||
|
}
|
||||||
|
for metric in (
|
||||||
|
"normalized_entropy",
|
||||||
|
"trajectory_span",
|
||||||
|
"gaussian_target_kl",
|
||||||
|
"mean_absolute_time_center_error",
|
||||||
|
):
|
||||||
|
raw = source.get(metric, "")
|
||||||
|
if raw not in (None, ""):
|
||||||
|
value = float(raw)
|
||||||
|
if np.isfinite(value):
|
||||||
|
row[metric] = value
|
||||||
|
grouped[(row["variant"], row["modality"])].append(row)
|
||||||
|
|
||||||
|
metric_names = (
|
||||||
|
"normalized_entropy",
|
||||||
|
"trajectory_span",
|
||||||
|
"gaussian_target_kl",
|
||||||
|
"mean_absolute_time_center_error",
|
||||||
|
)
|
||||||
|
results = []
|
||||||
|
for (variant, modality), rows in sorted(grouped.items()):
|
||||||
|
result: dict[str, Any] = {
|
||||||
|
"variant": variant,
|
||||||
|
"modality": modality,
|
||||||
|
"clip_count": len(rows),
|
||||||
|
"video_id_count": len({row["video_id"] for row in rows}),
|
||||||
|
}
|
||||||
|
for index, metric in enumerate(metric_names):
|
||||||
|
values = [row for row in rows if metric in row]
|
||||||
|
if not values:
|
||||||
|
continue
|
||||||
|
mean, low, high, _ = _cluster_bootstrap(
|
||||||
|
values,
|
||||||
|
metric,
|
||||||
|
seed=seed + index + sum(ord(character) for character in variant + modality),
|
||||||
|
)
|
||||||
|
result[f"{metric}_video_macro_mean"] = mean
|
||||||
|
result[f"{metric}_ci95_low"] = low
|
||||||
|
result[f"{metric}_ci95_high"] = high
|
||||||
|
results.append(result)
|
||||||
|
return results
|
||||||
|
|
||||||
|
|
||||||
|
def _plot_curve(curve_rows: Sequence[Mapping[str, Any]], output_path: Path) -> None:
|
||||||
|
pairs = tuple(f"{left}_{right}" for left, right in PAIRINGS)
|
||||||
|
colors = {
|
||||||
|
"M1": "#555555",
|
||||||
|
"M2": "#9a9a9a",
|
||||||
|
"M3_noSourceTime": "#2878b5",
|
||||||
|
"M3_sourceTime": "#f08a24",
|
||||||
|
"M4_noSourceTime": "#55a868",
|
||||||
|
"M4_sourceTime": "#c44e52",
|
||||||
|
}
|
||||||
|
fig, axes = plt.subplots(1, 3, figsize=(17, 5), sharey=True)
|
||||||
|
for axis, pair in zip(axes, pairs, strict=True):
|
||||||
|
for method in METHODS:
|
||||||
|
rows = [row for row in curve_rows if row["pair"] == pair and row["method"] == method]
|
||||||
|
if not rows:
|
||||||
|
continue
|
||||||
|
deltas = sorted({int(row["delta"]) for row in rows})
|
||||||
|
xs, means, lows, highs = [], [], [], []
|
||||||
|
for delta in deltas:
|
||||||
|
subset = [row for row in rows if int(row["delta"]) == delta]
|
||||||
|
mean, low, high, _ = _cluster_bootstrap(
|
||||||
|
subset,
|
||||||
|
"similarity",
|
||||||
|
seed=9821 + delta + sum(ord(character) for character in method + pair),
|
||||||
|
repetitions=1000,
|
||||||
|
)
|
||||||
|
xs.append(delta)
|
||||||
|
means.append(mean)
|
||||||
|
lows.append(low)
|
||||||
|
highs.append(high)
|
||||||
|
axis.plot(xs, means, marker="o", markersize=3, linewidth=1.6, label=method, color=colors[method])
|
||||||
|
axis.fill_between(xs, lows, highs, color=colors[method], alpha=0.10, linewidth=0)
|
||||||
|
axis.axvline(0, color="black", linestyle="--", linewidth=0.9, alpha=0.6)
|
||||||
|
axis.set_title(pair.replace("_", "–"))
|
||||||
|
axis.set_xlabel("Temporal shift Δ (slots)")
|
||||||
|
axis.grid(alpha=0.2)
|
||||||
|
axes[0].set_ylabel("Cross-modal cosine similarity")
|
||||||
|
handles, labels = axes[-1].get_legend_handles_labels()
|
||||||
|
fig.legend(handles, labels, loc="lower center", ncol=3, frameon=False, bbox_to_anchor=(0.5, -0.02))
|
||||||
|
fig.suptitle("Same-slot vs shifted cross-modal similarity (95% video_id bootstrap CI)", y=1.02)
|
||||||
|
fig.tight_layout(rect=(0, 0.08, 1, 0.98))
|
||||||
|
output_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
fig.savefig(output_path, dpi=180, bbox_inches="tight")
|
||||||
|
plt.close(fig)
|
||||||
|
|
||||||
|
|
||||||
|
def run(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
started = time.time()
|
||||||
|
if args.device == "auto":
|
||||||
|
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||||
|
else:
|
||||||
|
device = torch.device(args.device)
|
||||||
|
if device.type == "cuda" and not torch.cuda.is_available():
|
||||||
|
raise RuntimeError("CUDA was requested but is unavailable")
|
||||||
|
|
||||||
|
loaded_samples = load_feature_samples(args.feature_dir, args.manifest)
|
||||||
|
by_id = {sample.sample_id: sample for sample in loaded_samples}
|
||||||
|
split_rows = json.loads(args.splits.read_text(encoding="utf-8"))
|
||||||
|
output_dir = args.output_dir
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
all_metric_rows: list[dict[str, Any]] = []
|
||||||
|
all_curve_rows: list[dict[str, Any]] = []
|
||||||
|
all_history_rows: list[dict[str, Any]] = []
|
||||||
|
saved_probes: dict[str, Any] = {}
|
||||||
|
fold_manifests = []
|
||||||
|
|
||||||
|
for split in split_rows:
|
||||||
|
fold = int(split["fold"])
|
||||||
|
train_samples = [by_id[sample_id] for sample_id in split["train_sample_ids"]]
|
||||||
|
validation_samples = [by_id[sample_id] for sample_id in split["validation_sample_ids"]]
|
||||||
|
train_groups = {sample.group_id for sample in train_samples}
|
||||||
|
validation_groups = {sample.group_id for sample in validation_samples}
|
||||||
|
if train_groups & validation_groups:
|
||||||
|
raise ValueError(f"video_id leakage in fold {fold}")
|
||||||
|
stats = fit_feature_stats(train_samples)
|
||||||
|
print(
|
||||||
|
f"[correspondence fold {fold}] train={len(train_samples)} heldout={len(validation_samples)} "
|
||||||
|
f"video_ids={len(train_groups)}/{len(validation_groups)}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
fold_manifests.append(
|
||||||
|
{
|
||||||
|
"fold": fold,
|
||||||
|
"train_count": len(train_samples),
|
||||||
|
"heldout_count": len(validation_samples),
|
||||||
|
"train_video_ids": sorted(train_groups),
|
||||||
|
"heldout_video_ids": sorted(validation_groups),
|
||||||
|
"overlap": sorted(train_groups & validation_groups),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
for method_index, method_name in enumerate(METHODS):
|
||||||
|
if method_name in {"M1", "M2"}:
|
||||||
|
aligned_train, _ = _collect_representations(
|
||||||
|
method_name,
|
||||||
|
train_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
aligned_validation, _ = _collect_representations(
|
||||||
|
method_name,
|
||||||
|
validation_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
aligned = {**aligned_train, **aligned_validation}
|
||||||
|
else:
|
||||||
|
aligned = _collect_learned_representations(
|
||||||
|
method_name=method_name,
|
||||||
|
fold=fold,
|
||||||
|
train_samples=train_samples,
|
||||||
|
validation_samples=validation_samples,
|
||||||
|
stats=stats,
|
||||||
|
checkpoint_root=args.checkpoint_root,
|
||||||
|
device=device,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
probe_seed = args.seed + fold * 101 + method_index
|
||||||
|
probe, history = _fit_probe(
|
||||||
|
[sample.sample_id for sample in train_samples],
|
||||||
|
aligned,
|
||||||
|
device=device,
|
||||||
|
seed=probe_seed,
|
||||||
|
epochs=args.epochs,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
learning_rate=args.learning_rate,
|
||||||
|
temperature=args.temperature,
|
||||||
|
)
|
||||||
|
for row in history:
|
||||||
|
all_history_rows.append(
|
||||||
|
{"fold": fold, "method": method_name, "seed": probe_seed, **row}
|
||||||
|
)
|
||||||
|
probe.eval()
|
||||||
|
with torch.no_grad():
|
||||||
|
validation = _stack_ids(
|
||||||
|
[sample.sample_id for sample in validation_samples], aligned, device
|
||||||
|
)
|
||||||
|
projected = probe(validation)
|
||||||
|
for sample_index, sample in enumerate(validation_samples):
|
||||||
|
one = {name: projected[name][sample_index] for name in MODALITIES}
|
||||||
|
all_metric_rows.extend(
|
||||||
|
_sample_metrics(
|
||||||
|
method=method_name,
|
||||||
|
fold=fold,
|
||||||
|
sample=sample,
|
||||||
|
projected=one,
|
||||||
|
curve_rows=all_curve_rows,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
saved_probes[f"fold_{fold:02d}/{method_name}"] = {
|
||||||
|
"input_dimensions": {
|
||||||
|
name: int(aligned[train_samples[0].sample_id][name].shape[-1])
|
||||||
|
for name in MODALITIES
|
||||||
|
},
|
||||||
|
"state_dict": {key: value.detach().cpu() for key, value in probe.state_dict().items()},
|
||||||
|
"seed": probe_seed,
|
||||||
|
}
|
||||||
|
print(
|
||||||
|
f"[correspondence fold {fold} {method_name}] "
|
||||||
|
f"probe_loss={history[-1]['train_loss']:.4f}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
del probe
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
|
||||||
|
summary_rows, curve_summary_rows = _summaries(
|
||||||
|
all_metric_rows, all_curve_rows, seed=args.seed
|
||||||
|
)
|
||||||
|
time_localization_rows = _summarize_time_localization(
|
||||||
|
args.checkpoint_root, len(split_rows), seed=args.seed
|
||||||
|
)
|
||||||
|
_write_csv(output_dir / "correspondence_clip_metrics.csv", all_metric_rows)
|
||||||
|
_write_csv(output_dir / "shifted_similarity_by_clip.csv", all_curve_rows)
|
||||||
|
_write_csv(output_dir / "metric_summary.csv", summary_rows)
|
||||||
|
_write_csv(output_dir / "shift_curve_summary.csv", curve_summary_rows)
|
||||||
|
_write_csv(output_dir / "time_localization_summary.csv", time_localization_rows)
|
||||||
|
_write_csv(output_dir / "probe_training_history.csv", all_history_rows)
|
||||||
|
_plot_curve(all_curve_rows, output_dir / "shifted_similarity_curve.png")
|
||||||
|
torch.save(saved_probes, output_dir / "probe_checkpoints.pt")
|
||||||
|
manifest = {
|
||||||
|
"created_utc": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"experiment": "Q1-Correspondence Evaluation",
|
||||||
|
"sample_count": len(loaded_samples),
|
||||||
|
"fold_count": len(split_rows),
|
||||||
|
"folds": fold_manifests,
|
||||||
|
"methods": list(METHODS),
|
||||||
|
"grid_size": GRID_SIZE,
|
||||||
|
"probe": {
|
||||||
|
"type": "one bias-free linear projection per modality, L2 normalized",
|
||||||
|
"output_dimension": OUTPUT_SIZE,
|
||||||
|
"training_objective": "symmetric within-clip diagonal-vs-off-diagonal InfoNCE across all three modality pairs",
|
||||||
|
"negative_source": "same training clip, different grid slots",
|
||||||
|
"epochs": args.epochs,
|
||||||
|
"batch_size": args.batch_size,
|
||||||
|
"learning_rate": args.learning_rate,
|
||||||
|
"temperature": args.temperature,
|
||||||
|
"emotion_labels_used": False,
|
||||||
|
"heldout_groups_used_for_probe_training_or_selection": False,
|
||||||
|
},
|
||||||
|
"evaluation": {
|
||||||
|
"shift_curve": f"mean cosine sim(z_i^m, z_(i+delta)^n), delta=-{CURVE_MAX_SHIFT}..{CURVE_MAX_SHIFT}",
|
||||||
|
"temporal_retrieval": "candidate slots restricted to the same held-out clip; report exact R@1, within +/-1 R@1, and MASE slots in both directions",
|
||||||
|
"matched_vs_shifted_auc": f"per-clip ROC AUC; positives are same-slot pairs, negatives are same-clip pairs with |i-j|>{NEGATIVE_RADIUS}",
|
||||||
|
"confidence_intervals": "95% percentile bootstrap resampling video_id groups, not clips or slots; 2,000 repetitions for CSV summaries and 1,000 for plot ribbons",
|
||||||
|
"random_retrieval_reference": {
|
||||||
|
"exact_r1": 1.0 / GRID_SIZE,
|
||||||
|
"within_pm1_r1": (3.0 * GRID_SIZE - 2.0) / GRID_SIZE**2,
|
||||||
|
"expected_mase_slots": (GRID_SIZE**2 - 1) / (3.0 * GRID_SIZE),
|
||||||
|
},
|
||||||
|
},
|
||||||
|
"device": str(device),
|
||||||
|
"gpu_name": torch.cuda.get_device_name(device) if device.type == "cuda" else None,
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"elapsed_seconds": time.time() - started,
|
||||||
|
"interpretation_limits": [
|
||||||
|
"The projection is a supervised diagnostic probe for same-slot matchability, not an alignment model or independent ground truth.",
|
||||||
|
"A strong result means train-video same-time matchability transfers to held-out video_id groups; it does not establish semantic equivalence between modalities.",
|
||||||
|
"M3/M4 D5 training used timestamp-derived Gaussian targets; this evaluation tests whether their frozen pooled content representations carry transferable same-time correspondence beyond that temporal target.",
|
||||||
|
"The effective sample size for confidence intervals is the number of video_id groups, not the number of slots.",
|
||||||
|
],
|
||||||
|
}
|
||||||
|
(output_dir / "run_manifest.json").write_text(
|
||||||
|
json.dumps(manifest, ensure_ascii=False, indent=2, allow_nan=False), encoding="utf-8"
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f"[correspondence done] clips={len(all_metric_rows)} "
|
||||||
|
f"elapsed={manifest['elapsed_seconds']:.1f}s output={output_dir}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
return manifest
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
project = Path(__file__).resolve().parents[1]
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument("--device", choices=("auto", "cpu", "cuda"), default="auto")
|
||||||
|
parser.add_argument("--batch-size", type=int, default=8)
|
||||||
|
parser.add_argument("--epochs", type=int, default=40)
|
||||||
|
parser.add_argument("--seed", type=int, default=42)
|
||||||
|
parser.add_argument("--learning-rate", type=float, default=1e-3)
|
||||||
|
parser.add_argument("--temperature", type=float, default=0.1)
|
||||||
|
parser.add_argument(
|
||||||
|
"--feature-dir", type=Path, default=project / "outputs/q1_features/features"
|
||||||
|
)
|
||||||
|
parser.add_argument("--manifest", type=Path, default=project / "outputs/audit/manifest.csv")
|
||||||
|
parser.add_argument(
|
||||||
|
"--splits", type=Path, default=project / "outputs/method_comparison/splits.json"
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--checkpoint-root", type=Path, default=project / "outputs/alignment_debug/heldout"
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output-dir", type=Path, default=project / "outputs/correspondence_eval"
|
||||||
|
)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
run(args)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,159 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import csv
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Mapping, Sequence
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
from torch import Tensor
|
||||||
|
|
||||||
|
from .types import MODALITIES, SequenceBatch
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class FeatureSample:
|
||||||
|
sample_id: str
|
||||||
|
group_id: str
|
||||||
|
duration_s: float
|
||||||
|
sentiment: float
|
||||||
|
polarity: int
|
||||||
|
word_intervals: np.ndarray
|
||||||
|
features: Mapping[str, np.ndarray]
|
||||||
|
times: Mapping[str, np.ndarray]
|
||||||
|
valid: Mapping[str, np.ndarray]
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class FeatureStats:
|
||||||
|
mean: Mapping[str, np.ndarray]
|
||||||
|
scale: Mapping[str, np.ndarray]
|
||||||
|
|
||||||
|
|
||||||
|
def load_feature_samples(feature_dir: Path, manifest_path: Path) -> list[FeatureSample]:
|
||||||
|
"""Read the extracted NPZ files and their source-time/label manifest."""
|
||||||
|
with manifest_path.open("r", encoding="utf-8-sig", newline="") as file:
|
||||||
|
rows = list(csv.DictReader(file))
|
||||||
|
if not rows:
|
||||||
|
raise ValueError(f"sample manifest is empty: {manifest_path}")
|
||||||
|
|
||||||
|
samples: list[FeatureSample] = []
|
||||||
|
seen: set[str] = set()
|
||||||
|
for row in rows:
|
||||||
|
video_id = row["video_id"]
|
||||||
|
clip_id = row["clip_id"]
|
||||||
|
sample_id = f"{video_id}/{clip_id}"
|
||||||
|
if sample_id in seen:
|
||||||
|
raise ValueError(f"duplicate sample in manifest: {sample_id}")
|
||||||
|
seen.add(sample_id)
|
||||||
|
|
||||||
|
path = feature_dir / f"{video_id}__{clip_id}.npz"
|
||||||
|
if not path.is_file():
|
||||||
|
raise FileNotFoundError(f"feature file missing for {sample_id}: {path}")
|
||||||
|
with np.load(path, allow_pickle=False) as archive:
|
||||||
|
features = {
|
||||||
|
"text": np.asarray(archive["text_features"], dtype=np.float32),
|
||||||
|
"audio": np.asarray(archive["audio_features"], dtype=np.float32),
|
||||||
|
"vision": np.asarray(archive["vision_features"], dtype=np.float32),
|
||||||
|
}
|
||||||
|
times = {
|
||||||
|
"text": np.asarray(archive["word_intervals_s"], dtype=np.float32).mean(axis=1),
|
||||||
|
"audio": np.asarray(archive["audio_times_s"], dtype=np.float32),
|
||||||
|
"vision": np.asarray(archive["vision_times_s"], dtype=np.float32),
|
||||||
|
}
|
||||||
|
valid = {
|
||||||
|
"text": np.ones(features["text"].shape[0], dtype=np.bool_),
|
||||||
|
"audio": np.asarray(archive["audio_valid"], dtype=np.bool_),
|
||||||
|
"vision": np.asarray(archive["vision_valid"], dtype=np.bool_),
|
||||||
|
}
|
||||||
|
word_intervals = np.asarray(archive["word_intervals_s"], dtype=np.float32)
|
||||||
|
|
||||||
|
for name in MODALITIES:
|
||||||
|
if features[name].ndim != 2 or times[name].shape != (features[name].shape[0],):
|
||||||
|
raise ValueError(f"invalid {name} feature/timestamp shape in {sample_id}")
|
||||||
|
if valid[name].shape != times[name].shape or not valid[name].any():
|
||||||
|
raise ValueError(f"{sample_id} has no valid {name} sequence positions")
|
||||||
|
if not np.isfinite(features[name][valid[name]]).all():
|
||||||
|
raise ValueError(f"non-finite valid {name} values in {sample_id}")
|
||||||
|
if not np.isfinite(times[name][valid[name]]).all():
|
||||||
|
raise ValueError(f"non-finite valid {name} timestamps in {sample_id}")
|
||||||
|
|
||||||
|
annotation = row["annotation"].strip().lower()
|
||||||
|
polarity_by_name = {"negative": 0, "neutral": 1, "positive": 2}
|
||||||
|
if annotation not in polarity_by_name:
|
||||||
|
raise ValueError(f"unknown polarity label {annotation!r} in {sample_id}")
|
||||||
|
samples.append(
|
||||||
|
FeatureSample(
|
||||||
|
sample_id=sample_id,
|
||||||
|
group_id=row.get("group_id") or video_id,
|
||||||
|
duration_s=float(row["duration_s"]),
|
||||||
|
sentiment=float(row["label"]),
|
||||||
|
polarity=polarity_by_name[annotation],
|
||||||
|
word_intervals=word_intervals,
|
||||||
|
features=features,
|
||||||
|
times=times,
|
||||||
|
valid=valid,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return samples
|
||||||
|
|
||||||
|
|
||||||
|
def fit_feature_stats(samples: Sequence[FeatureSample]) -> FeatureStats:
|
||||||
|
"""Fit per-modality z-score parameters using only the training fold."""
|
||||||
|
if not samples:
|
||||||
|
raise ValueError("cannot fit feature statistics on an empty sample list")
|
||||||
|
means: dict[str, np.ndarray] = {}
|
||||||
|
scales: dict[str, np.ndarray] = {}
|
||||||
|
for name in MODALITIES:
|
||||||
|
values = np.concatenate(
|
||||||
|
[sample.features[name][sample.valid[name]] for sample in samples], axis=0
|
||||||
|
).astype(np.float64, copy=False)
|
||||||
|
mean = values.mean(axis=0)
|
||||||
|
scale = values.std(axis=0)
|
||||||
|
scale[scale < 1e-6] = 1.0
|
||||||
|
means[name] = mean.astype(np.float32)
|
||||||
|
scales[name] = scale.astype(np.float32)
|
||||||
|
return FeatureStats(mean=means, scale=scales)
|
||||||
|
|
||||||
|
|
||||||
|
def standardized_features(sample: FeatureSample, stats: FeatureStats) -> dict[str, np.ndarray]:
|
||||||
|
return {
|
||||||
|
name: ((sample.features[name] - stats.mean[name]) / stats.scale[name]).astype(
|
||||||
|
np.float32, copy=False
|
||||||
|
)
|
||||||
|
for name in MODALITIES
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def collate_feature_samples(
|
||||||
|
samples: Sequence[FeatureSample],
|
||||||
|
stats: FeatureStats,
|
||||||
|
device: torch.device,
|
||||||
|
) -> tuple[dict[str, SequenceBatch], Tensor, list[Tensor]]:
|
||||||
|
"""Pad one variable-length batch in memory; padding is masked and never saved."""
|
||||||
|
if not samples:
|
||||||
|
raise ValueError("cannot collate an empty sample list")
|
||||||
|
batch_size = len(samples)
|
||||||
|
sequences: dict[str, SequenceBatch] = {}
|
||||||
|
for name in MODALITIES:
|
||||||
|
lengths = [sample.features[name].shape[0] for sample in samples]
|
||||||
|
max_length = max(lengths)
|
||||||
|
dimension = samples[0].features[name].shape[1]
|
||||||
|
feature_batch = torch.zeros(batch_size, max_length, dimension, dtype=torch.float32)
|
||||||
|
time_batch = torch.zeros(batch_size, max_length, dtype=torch.float32)
|
||||||
|
valid_batch = torch.zeros(batch_size, max_length, dtype=torch.bool)
|
||||||
|
for index, sample in enumerate(samples):
|
||||||
|
features = standardized_features(sample, stats)[name]
|
||||||
|
length = len(features)
|
||||||
|
feature_batch[index, :length] = torch.from_numpy(features)
|
||||||
|
time_batch[index, :length] = torch.from_numpy(sample.times[name])
|
||||||
|
valid_batch[index, :length] = torch.from_numpy(sample.valid[name])
|
||||||
|
sequences[name] = SequenceBatch(
|
||||||
|
features=feature_batch.to(device),
|
||||||
|
times=time_batch.to(device),
|
||||||
|
valid=valid_batch.to(device),
|
||||||
|
)
|
||||||
|
durations = torch.tensor([sample.duration_s for sample in samples], dtype=torch.float32, device=device)
|
||||||
|
word_intervals = [torch.from_numpy(sample.word_intervals).to(device) for sample in samples]
|
||||||
|
return sequences, durations, word_intervals
|
||||||
@@ -0,0 +1,462 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from collections.abc import Mapping, Sequence
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
import torch.nn.functional as F
|
||||||
|
from sklearn.dummy import DummyClassifier
|
||||||
|
from sklearn.linear_model import LogisticRegression, Ridge
|
||||||
|
from sklearn.metrics import accuracy_score, f1_score, mean_absolute_error
|
||||||
|
from sklearn.preprocessing import StandardScaler
|
||||||
|
from torch import Tensor, nn
|
||||||
|
|
||||||
|
from .alignment import make_block_mask
|
||||||
|
from .losses import cross_modal_contrastive_loss
|
||||||
|
from .metrics import retrieval_metrics
|
||||||
|
from .types import MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
class RetrievalProjection(nn.Module):
|
||||||
|
"""Equal-capacity linear heads for the cross-modal retrieval probe."""
|
||||||
|
|
||||||
|
def __init__(self, dimensions: Mapping[str, int], output_size: int = 128) -> None:
|
||||||
|
super().__init__()
|
||||||
|
self.projections = nn.ModuleDict(
|
||||||
|
{name: nn.Linear(dimensions[name], output_size, bias=False) for name in MODALITIES}
|
||||||
|
)
|
||||||
|
|
||||||
|
def forward(self, aligned: Mapping[str, Tensor]) -> dict[str, Tensor]:
|
||||||
|
return {
|
||||||
|
name: F.normalize(self.projections[name](aligned[name]), dim=-1)
|
||||||
|
for name in MODALITIES
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class MaskedCrossModalDecoder(nn.Module):
|
||||||
|
"""Reconstruct one missing modality from the other aligned streams."""
|
||||||
|
|
||||||
|
def __init__(self, dimensions: Mapping[str, int], hidden_size: int = 256) -> None:
|
||||||
|
super().__init__()
|
||||||
|
self.dimensions = dict(dimensions)
|
||||||
|
input_size = sum(dimensions.values()) + len(MODALITIES)
|
||||||
|
self.decoders = nn.ModuleDict(
|
||||||
|
{
|
||||||
|
target: nn.Sequential(
|
||||||
|
nn.Linear(input_size, hidden_size),
|
||||||
|
nn.GELU(),
|
||||||
|
nn.Dropout(0.1),
|
||||||
|
nn.Linear(hidden_size, dimensions[target]),
|
||||||
|
)
|
||||||
|
for target in MODALITIES
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
def forward(self, aligned: Mapping[str, Tensor], target: str, mask: Tensor) -> Tensor:
|
||||||
|
values = []
|
||||||
|
availability = []
|
||||||
|
for name in MODALITIES:
|
||||||
|
present = torch.ones_like(mask, dtype=aligned[name].dtype)
|
||||||
|
source = aligned[name]
|
||||||
|
if name == target:
|
||||||
|
present = (~mask).to(source.dtype)
|
||||||
|
source = source.masked_fill(mask.unsqueeze(-1), 0.0)
|
||||||
|
values.append(source)
|
||||||
|
availability.append(present.unsqueeze(-1))
|
||||||
|
inputs = torch.cat((*values, *availability), dim=-1)
|
||||||
|
return self.decoders[target](inputs)
|
||||||
|
|
||||||
|
|
||||||
|
def _stack_aligned(
|
||||||
|
sample_ids: Sequence[str], aligned_by_id: Mapping[str, Mapping[str, np.ndarray]], device: torch.device
|
||||||
|
) -> dict[str, Tensor]:
|
||||||
|
return {
|
||||||
|
name: torch.as_tensor(
|
||||||
|
np.stack([aligned_by_id[sample_id][name] for sample_id in sample_ids]),
|
||||||
|
dtype=torch.float32,
|
||||||
|
device=device,
|
||||||
|
)
|
||||||
|
for name in MODALITIES
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def run_retrieval_probe(
|
||||||
|
train_ids: Sequence[str],
|
||||||
|
val_ids: Sequence[str],
|
||||||
|
aligned_by_id: Mapping[str, Mapping[str, np.ndarray]],
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
seed: int,
|
||||||
|
epochs: int = 20,
|
||||||
|
batch_size: int = 16,
|
||||||
|
learning_rate: float = 1e-3,
|
||||||
|
temperature: float = 0.1,
|
||||||
|
) -> list[dict[str, float | str]]:
|
||||||
|
"""Fit modality projections on training clips, then score held-out retrieval.
|
||||||
|
|
||||||
|
The positive is a matching common-grid index within a clip. This measures
|
||||||
|
representation consistency; it is not independent temporal ground truth.
|
||||||
|
"""
|
||||||
|
if not train_ids or not val_ids:
|
||||||
|
raise ValueError("retrieval probe needs non-empty train and validation sets")
|
||||||
|
device_gen = torch.Generator(device=device)
|
||||||
|
device_gen.manual_seed(seed)
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
train = _stack_aligned(train_ids, aligned_by_id, device)
|
||||||
|
validation = _stack_aligned(val_ids, aligned_by_id, device)
|
||||||
|
dimensions = {name: int(train[name].shape[-1]) for name in MODALITIES}
|
||||||
|
model = RetrievalProjection(dimensions).to(device)
|
||||||
|
optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate, weight_decay=1e-4)
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
model.train()
|
||||||
|
for _ in range(epochs):
|
||||||
|
order = rng.permutation(len(train_ids))
|
||||||
|
for start in range(0, len(order), batch_size):
|
||||||
|
indices = torch.as_tensor(order[start : start + batch_size], device=device)
|
||||||
|
batch = {name: value.index_select(0, indices) for name, value in train.items()}
|
||||||
|
projected = model(batch)
|
||||||
|
loss = cross_modal_contrastive_loss(projected, temperature=temperature)
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
loss.backward()
|
||||||
|
nn.utils.clip_grad_norm_(model.parameters(), 1.0)
|
||||||
|
optimizer.step()
|
||||||
|
|
||||||
|
model.eval()
|
||||||
|
with torch.no_grad():
|
||||||
|
projected = model(validation)
|
||||||
|
directions = (("text", "audio"), ("audio", "text"), ("text", "vision"),
|
||||||
|
("vision", "text"), ("audio", "vision"), ("vision", "audio"))
|
||||||
|
rows: list[dict[str, float | str]] = []
|
||||||
|
for query_name, target_name in directions:
|
||||||
|
values = retrieval_metrics(projected[query_name], projected[target_name])
|
||||||
|
rows.append({"direction": f"{query_name}_to_{target_name}", **values})
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def run_within_clip_temporal_retrieval_probe(
|
||||||
|
train_ids: Sequence[str],
|
||||||
|
val_ids: Sequence[str],
|
||||||
|
aligned_by_id: Mapping[str, Mapping[str, np.ndarray]],
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
seed: int,
|
||||||
|
epochs: int = 20,
|
||||||
|
batch_size: int = 16,
|
||||||
|
learning_rate: float = 1e-3,
|
||||||
|
temperature: float = 0.1,
|
||||||
|
tolerance: int = 1,
|
||||||
|
top_k: int = 3,
|
||||||
|
) -> list[dict[str, float | str]]:
|
||||||
|
"""Fit train-only cross-modal projections, then retrieve slots within each clip.
|
||||||
|
|
||||||
|
Unlike global grid retrieval, each query's candidates are restricted to
|
||||||
|
the target modality from that same held-out clip. A result is correct if
|
||||||
|
its slot is within ``tolerance`` of the query slot.
|
||||||
|
"""
|
||||||
|
if not train_ids or not val_ids:
|
||||||
|
raise ValueError("temporal retrieval needs non-empty train and validation sets")
|
||||||
|
if tolerance < 0 or top_k < 1:
|
||||||
|
raise ValueError("tolerance must be non-negative and top_k positive")
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
train = _stack_aligned(train_ids, aligned_by_id, device)
|
||||||
|
validation = _stack_aligned(val_ids, aligned_by_id, device)
|
||||||
|
dimensions = {name: int(train[name].shape[-1]) for name in MODALITIES}
|
||||||
|
model = RetrievalProjection(dimensions).to(device)
|
||||||
|
optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate, weight_decay=1e-4)
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
model.train()
|
||||||
|
for _ in range(epochs):
|
||||||
|
order = rng.permutation(len(train_ids))
|
||||||
|
for start in range(0, len(order), batch_size):
|
||||||
|
indices = torch.as_tensor(order[start : start + batch_size], device=device)
|
||||||
|
batch = {name: value.index_select(0, indices) for name, value in train.items()}
|
||||||
|
projected = model(batch)
|
||||||
|
loss = cross_modal_contrastive_loss(projected, temperature=temperature)
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
loss.backward()
|
||||||
|
nn.utils.clip_grad_norm_(model.parameters(), 1.0)
|
||||||
|
optimizer.step()
|
||||||
|
|
||||||
|
model.eval()
|
||||||
|
with torch.no_grad():
|
||||||
|
projected = model(validation)
|
||||||
|
directions = (
|
||||||
|
("text", "audio"),
|
||||||
|
("audio", "text"),
|
||||||
|
("text", "vision"),
|
||||||
|
("vision", "text"),
|
||||||
|
("audio", "vision"),
|
||||||
|
("vision", "audio"),
|
||||||
|
)
|
||||||
|
rows: list[dict[str, float | str]] = []
|
||||||
|
for query_name, target_name in directions:
|
||||||
|
distances_top1 = []
|
||||||
|
hits_top1 = []
|
||||||
|
hits_topk = []
|
||||||
|
for clip_index in range(len(val_ids)):
|
||||||
|
query = projected[query_name][clip_index]
|
||||||
|
target = projected[target_name][clip_index]
|
||||||
|
scores = query @ target.T
|
||||||
|
count = scores.shape[0]
|
||||||
|
k = min(top_k, count)
|
||||||
|
candidates = scores.topk(k=k, dim=-1).indices
|
||||||
|
slots = torch.arange(count, device=device)[:, None]
|
||||||
|
distances = (candidates - slots).abs()
|
||||||
|
distances_top1.append(distances[:, 0].float())
|
||||||
|
hits_top1.append((distances[:, 0] <= tolerance).float())
|
||||||
|
hits_topk.append((distances <= tolerance).any(dim=-1).float())
|
||||||
|
top1_distance = torch.cat(distances_top1)
|
||||||
|
rows.append(
|
||||||
|
{
|
||||||
|
"direction": f"{query_name}_to_{target_name}",
|
||||||
|
"r_at_1": float(torch.cat(hits_top1).mean().item()),
|
||||||
|
"r_at_3": float(torch.cat(hits_topk).mean().item()),
|
||||||
|
"mase_slots": float(top1_distance.mean().item()),
|
||||||
|
"exact_r_at_1": float((top1_distance == 0).float().mean().item()),
|
||||||
|
"tolerance_slots": float(tolerance),
|
||||||
|
"candidate_slots_per_clip": float(projected[query_name].shape[1]),
|
||||||
|
"queries": float(len(val_ids) * projected[query_name].shape[1]),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def _fixed_block_mask(
|
||||||
|
count: int, grid_size: int, ratio: float, device: torch.device, salt: int
|
||||||
|
) -> Tensor:
|
||||||
|
block = min(max(1, round(grid_size * ratio)), grid_size - 1)
|
||||||
|
starts = torch.tensor(
|
||||||
|
[(index * 17 + salt * 13) % (grid_size - block + 1) for index in range(count)],
|
||||||
|
dtype=torch.long,
|
||||||
|
device=device,
|
||||||
|
)
|
||||||
|
offsets = torch.arange(block, device=device)
|
||||||
|
mask = torch.zeros(count, grid_size, dtype=torch.bool, device=device)
|
||||||
|
mask[torch.arange(count, device=device)[:, None], starts[:, None] + offsets] = True
|
||||||
|
return mask
|
||||||
|
|
||||||
|
|
||||||
|
def run_reconstruction_probe(
|
||||||
|
train_ids: Sequence[str],
|
||||||
|
val_ids: Sequence[str],
|
||||||
|
aligned_by_id: Mapping[str, Mapping[str, np.ndarray]],
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
seed: int,
|
||||||
|
ratio: float = 0.2,
|
||||||
|
epochs: int = 25,
|
||||||
|
batch_size: int = 16,
|
||||||
|
learning_rate: float = 1e-3,
|
||||||
|
) -> list[dict[str, float | str]]:
|
||||||
|
"""Train the same decoder family on frozen alignments and score held-out clips."""
|
||||||
|
if not train_ids or not val_ids:
|
||||||
|
raise ValueError("reconstruction probe needs non-empty train and validation sets")
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
generator = torch.Generator(device=device)
|
||||||
|
generator.manual_seed(seed)
|
||||||
|
train = _stack_aligned(train_ids, aligned_by_id, device)
|
||||||
|
validation = _stack_aligned(val_ids, aligned_by_id, device)
|
||||||
|
dimensions = {name: int(train[name].shape[-1]) for name in MODALITIES}
|
||||||
|
model = MaskedCrossModalDecoder(dimensions).to(device)
|
||||||
|
optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate, weight_decay=1e-4)
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
|
||||||
|
model.train()
|
||||||
|
for _ in range(epochs):
|
||||||
|
order = rng.permutation(len(train_ids))
|
||||||
|
for start in range(0, len(order), batch_size):
|
||||||
|
indices = torch.as_tensor(order[start : start + batch_size], device=device)
|
||||||
|
batch = {name: value.index_select(0, indices) for name, value in train.items()}
|
||||||
|
target_losses = []
|
||||||
|
for target in MODALITIES:
|
||||||
|
mask = make_block_mask(
|
||||||
|
len(indices),
|
||||||
|
batch[target].shape[1],
|
||||||
|
ratio,
|
||||||
|
device,
|
||||||
|
generator=generator,
|
||||||
|
)
|
||||||
|
prediction = model(batch, target, mask)
|
||||||
|
target_losses.append(F.smooth_l1_loss(prediction[mask], batch[target][mask]))
|
||||||
|
loss = torch.stack(target_losses).mean()
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
loss.backward()
|
||||||
|
nn.utils.clip_grad_norm_(model.parameters(), 1.0)
|
||||||
|
optimizer.step()
|
||||||
|
|
||||||
|
model.eval()
|
||||||
|
rows: list[dict[str, float | str]] = []
|
||||||
|
with torch.no_grad():
|
||||||
|
for target_index, target in enumerate(MODALITIES):
|
||||||
|
mask = _fixed_block_mask(
|
||||||
|
len(val_ids), validation[target].shape[1], ratio, device, target_index
|
||||||
|
)
|
||||||
|
prediction = model(validation, target, mask)
|
||||||
|
residual = (prediction[mask] - validation[target][mask]).abs()
|
||||||
|
smooth = F.smooth_l1_loss(prediction[mask], validation[target][mask])
|
||||||
|
rows.append(
|
||||||
|
{
|
||||||
|
"target_modality": target,
|
||||||
|
"mask_ratio": ratio,
|
||||||
|
"mae_standardized": float(residual.mean().item()),
|
||||||
|
"smooth_l1_standardized": float(smooth.item()),
|
||||||
|
"masked_values": int(residual.numel()),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def run_shuffled_alignment_reconstruction_probe(
|
||||||
|
train_ids: Sequence[str],
|
||||||
|
val_ids: Sequence[str],
|
||||||
|
aligned_by_id: Mapping[str, Mapping[str, np.ndarray]],
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
seed: int,
|
||||||
|
ratio: float = 0.2,
|
||||||
|
epochs: int = 25,
|
||||||
|
batch_size: int = 16,
|
||||||
|
learning_rate: float = 1e-3,
|
||||||
|
shuffle_repeats: int = 5,
|
||||||
|
) -> list[dict[str, float | str]]:
|
||||||
|
"""Compare normal reconstruction with cross-modal slots shuffled at eval.
|
||||||
|
|
||||||
|
The decoder is trained once on aligned training-fold representations.
|
||||||
|
For the control, the two non-target modalities share a random slot
|
||||||
|
permutation within each validation clip; the target stream and target
|
||||||
|
values remain in their original order. Thus the metric isolates how much
|
||||||
|
correctly matched cross-modal slots help the frozen decoder.
|
||||||
|
"""
|
||||||
|
if not train_ids or not val_ids:
|
||||||
|
raise ValueError("reconstruction needs non-empty train and validation sets")
|
||||||
|
if shuffle_repeats < 1:
|
||||||
|
raise ValueError("shuffle_repeats must be positive")
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
generator = torch.Generator(device=device)
|
||||||
|
generator.manual_seed(seed)
|
||||||
|
train = _stack_aligned(train_ids, aligned_by_id, device)
|
||||||
|
validation = _stack_aligned(val_ids, aligned_by_id, device)
|
||||||
|
dimensions = {name: int(train[name].shape[-1]) for name in MODALITIES}
|
||||||
|
model = MaskedCrossModalDecoder(dimensions).to(device)
|
||||||
|
optimizer = torch.optim.AdamW(model.parameters(), lr=learning_rate, weight_decay=1e-4)
|
||||||
|
rng = np.random.default_rng(seed)
|
||||||
|
|
||||||
|
model.train()
|
||||||
|
for _ in range(epochs):
|
||||||
|
order = rng.permutation(len(train_ids))
|
||||||
|
for start in range(0, len(order), batch_size):
|
||||||
|
indices = torch.as_tensor(order[start : start + batch_size], device=device)
|
||||||
|
batch = {name: value.index_select(0, indices) for name, value in train.items()}
|
||||||
|
target_losses = []
|
||||||
|
for target in MODALITIES:
|
||||||
|
mask = make_block_mask(
|
||||||
|
len(indices), batch[target].shape[1], ratio, device, generator=generator
|
||||||
|
)
|
||||||
|
prediction = model(batch, target, mask)
|
||||||
|
target_losses.append(F.smooth_l1_loss(prediction[mask], batch[target][mask]))
|
||||||
|
loss = torch.stack(target_losses).mean()
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
loss.backward()
|
||||||
|
nn.utils.clip_grad_norm_(model.parameters(), 1.0)
|
||||||
|
optimizer.step()
|
||||||
|
|
||||||
|
model.eval()
|
||||||
|
rows: list[dict[str, float | str]] = []
|
||||||
|
with torch.no_grad():
|
||||||
|
for target_index, target in enumerate(MODALITIES):
|
||||||
|
mask = _fixed_block_mask(
|
||||||
|
len(val_ids), validation[target].shape[1], ratio, device, target_index
|
||||||
|
)
|
||||||
|
aligned_prediction = model(validation, target, mask)
|
||||||
|
aligned_error = (aligned_prediction[mask] - validation[target][mask]).abs().mean()
|
||||||
|
|
||||||
|
shuffle_errors = []
|
||||||
|
permutation_rng = np.random.default_rng(seed + 1709 + target_index)
|
||||||
|
slot_count = validation[target].shape[1]
|
||||||
|
for _ in range(shuffle_repeats):
|
||||||
|
permutations = np.stack(
|
||||||
|
[permutation_rng.permutation(slot_count) for _ in val_ids]
|
||||||
|
)
|
||||||
|
permutation_tensor = torch.as_tensor(permutations, dtype=torch.long, device=device)
|
||||||
|
shuffled = {}
|
||||||
|
for name, values in validation.items():
|
||||||
|
if name == target:
|
||||||
|
shuffled[name] = values
|
||||||
|
else:
|
||||||
|
gather_indices = permutation_tensor.unsqueeze(-1).expand_as(values)
|
||||||
|
shuffled[name] = values.gather(1, gather_indices)
|
||||||
|
shuffled_prediction = model(shuffled, target, mask)
|
||||||
|
shuffled_error = (
|
||||||
|
shuffled_prediction[mask] - validation[target][mask]
|
||||||
|
).abs().mean()
|
||||||
|
shuffle_errors.append(float(shuffled_error.item()))
|
||||||
|
shuffled_mean = float(np.mean(shuffle_errors))
|
||||||
|
rows.append(
|
||||||
|
{
|
||||||
|
"target_modality": target,
|
||||||
|
"mask_ratio": ratio,
|
||||||
|
"mae_aligned": float(aligned_error.item()),
|
||||||
|
"mae_shuffled_mean": shuffled_mean,
|
||||||
|
"mae_shuffled_std": float(np.std(shuffle_errors, ddof=1))
|
||||||
|
if shuffle_repeats > 1
|
||||||
|
else 0.0,
|
||||||
|
"gain_align": shuffled_mean - float(aligned_error.item()),
|
||||||
|
"shuffle_repeats": float(shuffle_repeats),
|
||||||
|
"masked_values": int(mask.sum().item() * dimensions[target]),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def _emotion_features(
|
||||||
|
sample_ids: Sequence[str], aligned_by_id: Mapping[str, Mapping[str, np.ndarray]], bins: int = 5
|
||||||
|
) -> np.ndarray:
|
||||||
|
outputs = []
|
||||||
|
for sample_id in sample_ids:
|
||||||
|
combined = np.concatenate([aligned_by_id[sample_id][name] for name in MODALITIES], axis=-1)
|
||||||
|
segments = np.array_split(combined, bins, axis=0)
|
||||||
|
outputs.append(np.concatenate([segment.mean(axis=0) for segment in segments]))
|
||||||
|
return np.stack(outputs).astype(np.float32, copy=False)
|
||||||
|
|
||||||
|
|
||||||
|
def run_frozen_emotion_probe(
|
||||||
|
train_samples: Sequence[object],
|
||||||
|
val_samples: Sequence[object],
|
||||||
|
aligned_by_id: Mapping[str, Mapping[str, np.ndarray]],
|
||||||
|
) -> dict[str, float | int]:
|
||||||
|
"""Evaluate a regularized five-bin linear probe on frozen aligned features."""
|
||||||
|
train_ids = [sample.sample_id for sample in train_samples]
|
||||||
|
val_ids = [sample.sample_id for sample in val_samples]
|
||||||
|
x_train = _emotion_features(train_ids, aligned_by_id)
|
||||||
|
x_val = _emotion_features(val_ids, aligned_by_id)
|
||||||
|
y_train = np.asarray([sample.polarity for sample in train_samples], dtype=np.int64)
|
||||||
|
y_val = np.asarray([sample.polarity for sample in val_samples], dtype=np.int64)
|
||||||
|
target_train = np.asarray([sample.sentiment for sample in train_samples], dtype=np.float64)
|
||||||
|
target_val = np.asarray([sample.sentiment for sample in val_samples], dtype=np.float64)
|
||||||
|
|
||||||
|
scaler = StandardScaler()
|
||||||
|
x_train = scaler.fit_transform(x_train)
|
||||||
|
x_val = scaler.transform(x_val)
|
||||||
|
if np.unique(y_train).size > 1:
|
||||||
|
classifier = LogisticRegression(
|
||||||
|
C=0.1, class_weight="balanced", max_iter=2000, solver="lbfgs", random_state=0
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
classifier = DummyClassifier(strategy="most_frequent")
|
||||||
|
classifier.fit(x_train, y_train)
|
||||||
|
predicted_class = classifier.predict(x_val)
|
||||||
|
regressor = Ridge(alpha=10.0)
|
||||||
|
regressor.fit(x_train, target_train)
|
||||||
|
predicted_score = regressor.predict(x_val)
|
||||||
|
if len(target_val) > 1 and np.std(predicted_score) > 0 and np.std(target_val) > 0:
|
||||||
|
pearson = float(np.corrcoef(predicted_score, target_val)[0, 1])
|
||||||
|
else:
|
||||||
|
pearson = float("nan")
|
||||||
|
return {
|
||||||
|
"accuracy": float(accuracy_score(y_val, predicted_class)),
|
||||||
|
"macro_f1": float(f1_score(y_val, predicted_class, average="macro", zero_division=0)),
|
||||||
|
"mae": float(mean_absolute_error(target_val, predicted_score)),
|
||||||
|
"pearson": pearson,
|
||||||
|
"n_train": len(train_samples),
|
||||||
|
"n_validation": len(val_samples),
|
||||||
|
}
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,175 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from collections.abc import Mapping
|
||||||
|
|
||||||
|
import torch
|
||||||
|
import torch.nn.functional as F
|
||||||
|
from torch import Tensor
|
||||||
|
|
||||||
|
from .metrics import attention_row_similarity
|
||||||
|
from .types import AlignmentOutput, MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
def temporal_monotonicity_loss(
|
||||||
|
output: AlignmentOutput,
|
||||||
|
times: Mapping[str, Tensor],
|
||||||
|
durations: Tensor,
|
||||||
|
epsilon: float = 0.02,
|
||||||
|
) -> Tensor:
|
||||||
|
"""Penalize backward motion on the normalized clip timeline."""
|
||||||
|
if epsilon < 0:
|
||||||
|
raise ValueError("epsilon must be non-negative")
|
||||||
|
losses = []
|
||||||
|
for name in MODALITIES:
|
||||||
|
mu = torch.bmm(output.weights[name], times[name].unsqueeze(-1)).squeeze(-1)
|
||||||
|
mu = mu / durations[:, None].clamp_min(torch.finfo(mu.dtype).eps)
|
||||||
|
backward = F.relu(mu[:, :-1] - mu[:, 1:] - epsilon)
|
||||||
|
losses.append(backward.square().mean())
|
||||||
|
return torch.stack(losses).mean()
|
||||||
|
|
||||||
|
|
||||||
|
def cross_modal_contrastive_loss(
|
||||||
|
aligned: Mapping[str, Tensor], temperature: float = 0.1
|
||||||
|
) -> Tensor:
|
||||||
|
"""Symmetric in-batch InfoNCE over same-sample, same-grid-slot positives."""
|
||||||
|
if temperature <= 0:
|
||||||
|
raise ValueError("temperature must be positive")
|
||||||
|
if set(aligned) != set(MODALITIES):
|
||||||
|
raise ValueError(f"aligned must contain exactly {MODALITIES}")
|
||||||
|
pair_losses = []
|
||||||
|
for left_index, left_name in enumerate(MODALITIES):
|
||||||
|
for right_name in MODALITIES[left_index + 1 :]:
|
||||||
|
left = F.normalize(aligned[left_name].flatten(0, 1), dim=-1)
|
||||||
|
right = F.normalize(aligned[right_name].flatten(0, 1), dim=-1)
|
||||||
|
if left.shape != right.shape:
|
||||||
|
raise ValueError("contrastive representations must share [B, K, D]")
|
||||||
|
logits = left @ right.T / temperature
|
||||||
|
labels = torch.arange(logits.shape[0], device=logits.device)
|
||||||
|
pair_losses.append(
|
||||||
|
(F.cross_entropy(logits, labels) + F.cross_entropy(logits.T, labels)) / 2
|
||||||
|
)
|
||||||
|
return torch.stack(pair_losses).mean()
|
||||||
|
|
||||||
|
|
||||||
|
def masked_reconstruction_loss(prediction: Tensor, target: Tensor, mask: Tensor) -> Tensor:
|
||||||
|
"""Smooth-L1 loss over masked grid slots only."""
|
||||||
|
if prediction.shape != target.shape:
|
||||||
|
raise ValueError("prediction and target must have the same shape")
|
||||||
|
if mask.shape != target.shape[:2] or mask.dtype != torch.bool:
|
||||||
|
raise ValueError("mask must be boolean with shape [B, K]")
|
||||||
|
if not bool(mask.any()):
|
||||||
|
raise ValueError("mask must select at least one target slot")
|
||||||
|
element_loss = F.smooth_l1_loss(prediction, target, reduction="none")
|
||||||
|
return element_loss[mask].mean()
|
||||||
|
|
||||||
|
|
||||||
|
def temporal_span_loss(
|
||||||
|
output: AlignmentOutput,
|
||||||
|
times: Mapping[str, Tensor],
|
||||||
|
durations: Tensor,
|
||||||
|
minimum_span: float = 0.7,
|
||||||
|
modalities: tuple[str, ...] = MODALITIES,
|
||||||
|
) -> Tensor:
|
||||||
|
"""Penalize an expected-time path that does not cover enough of a clip."""
|
||||||
|
if not 0.0 <= minimum_span <= 1.0:
|
||||||
|
raise ValueError("minimum_span must be in [0, 1]")
|
||||||
|
if not modalities or any(name not in MODALITIES for name in modalities):
|
||||||
|
raise ValueError("modalities must be a non-empty subset of MODALITIES")
|
||||||
|
losses = []
|
||||||
|
for name in modalities:
|
||||||
|
mu = torch.bmm(output.weights[name], times[name].unsqueeze(-1)).squeeze(-1)
|
||||||
|
mu = mu / durations[:, None].clamp_min(torch.finfo(mu.dtype).eps)
|
||||||
|
span = mu[:, -1] - mu[:, 0]
|
||||||
|
losses.append(F.relu(minimum_span - span).square().mean())
|
||||||
|
return torch.stack(losses).mean()
|
||||||
|
|
||||||
|
|
||||||
|
def attention_diversity_loss(
|
||||||
|
output: AlignmentOutput,
|
||||||
|
modalities: tuple[str, ...] = MODALITIES,
|
||||||
|
min_separation: int = 6,
|
||||||
|
) -> Tensor:
|
||||||
|
"""Penalize similar attention rows for slots far apart on the grid."""
|
||||||
|
if min_separation < 1:
|
||||||
|
raise ValueError("min_separation must be at least one")
|
||||||
|
if not modalities or any(name not in MODALITIES for name in modalities):
|
||||||
|
raise ValueError("modalities must be a non-empty subset of MODALITIES")
|
||||||
|
return torch.stack(
|
||||||
|
[attention_row_similarity(output.weights[name], min_separation).mean() for name in modalities]
|
||||||
|
).mean()
|
||||||
|
|
||||||
|
|
||||||
|
def weak_temporal_band_loss(
|
||||||
|
output: AlignmentOutput,
|
||||||
|
times: Mapping[str, Tensor],
|
||||||
|
durations: Tensor,
|
||||||
|
targets: Mapping[str, Tensor],
|
||||||
|
margin: float = 0.1,
|
||||||
|
) -> Tensor:
|
||||||
|
"""Allow soft alignment while keeping expected times near weak slot anchors."""
|
||||||
|
if margin < 0:
|
||||||
|
raise ValueError("margin must be non-negative")
|
||||||
|
if not targets:
|
||||||
|
return next(iter(output.weights.values())).sum() * 0.0
|
||||||
|
losses = []
|
||||||
|
for name, target in targets.items():
|
||||||
|
if name not in MODALITIES:
|
||||||
|
raise ValueError(f"unknown modality in temporal targets: {name}")
|
||||||
|
mu = torch.bmm(output.weights[name], times[name].unsqueeze(-1)).squeeze(-1)
|
||||||
|
mu = mu / durations[:, None].clamp_min(torch.finfo(mu.dtype).eps)
|
||||||
|
if target.shape != mu.shape:
|
||||||
|
raise ValueError(f"band target for {name} must have shape {tuple(mu.shape)}")
|
||||||
|
distance = (mu - target).abs()
|
||||||
|
losses.append(F.relu(distance - margin).square().mean())
|
||||||
|
return torch.stack(losses).mean()
|
||||||
|
|
||||||
|
|
||||||
|
def alignment_training_loss(
|
||||||
|
output: AlignmentOutput,
|
||||||
|
times: Mapping[str, Tensor],
|
||||||
|
durations: Tensor,
|
||||||
|
reconstruction: Tensor,
|
||||||
|
*,
|
||||||
|
lambda_rec: float = 1.0,
|
||||||
|
lambda_con: float = 1.0,
|
||||||
|
lambda_mono: float = 0.1,
|
||||||
|
lambda_span: float = 0.0,
|
||||||
|
lambda_div: float = 0.0,
|
||||||
|
lambda_band: float = 0.0,
|
||||||
|
epsilon: float = 0.02,
|
||||||
|
minimum_span: float = 0.7,
|
||||||
|
coverage_modalities: tuple[str, ...] = MODALITIES,
|
||||||
|
diversity_modalities: tuple[str, ...] = MODALITIES,
|
||||||
|
diversity_min_separation: int = 6,
|
||||||
|
band_targets: Mapping[str, Tensor] | None = None,
|
||||||
|
band_margin: float = 0.1,
|
||||||
|
) -> tuple[Tensor, dict[str, Tensor]]:
|
||||||
|
"""Shared M3/M4 training objective; emotion labels are deliberately unused."""
|
||||||
|
contrastive = cross_modal_contrastive_loss(output.aligned)
|
||||||
|
monotonicity = temporal_monotonicity_loss(output, times, durations, epsilon)
|
||||||
|
span = temporal_span_loss(
|
||||||
|
output, times, durations, minimum_span, modalities=coverage_modalities
|
||||||
|
)
|
||||||
|
diversity = attention_diversity_loss(
|
||||||
|
output, diversity_modalities, min_separation=diversity_min_separation
|
||||||
|
)
|
||||||
|
band = weak_temporal_band_loss(
|
||||||
|
output, times, durations, band_targets or {}, margin=band_margin
|
||||||
|
)
|
||||||
|
total = (
|
||||||
|
lambda_rec * reconstruction
|
||||||
|
+ lambda_con * contrastive
|
||||||
|
+ lambda_mono * monotonicity
|
||||||
|
+ lambda_span * span
|
||||||
|
+ lambda_div * diversity
|
||||||
|
+ lambda_band * band
|
||||||
|
)
|
||||||
|
return total, {
|
||||||
|
"reconstruction": reconstruction,
|
||||||
|
"contrastive": contrastive,
|
||||||
|
"monotonicity": monotonicity,
|
||||||
|
"span": span,
|
||||||
|
"diversity": diversity,
|
||||||
|
"band": band,
|
||||||
|
"total": total,
|
||||||
|
}
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,146 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
import torch.nn.functional as F
|
||||||
|
from torch import Tensor
|
||||||
|
|
||||||
|
from .types import MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
def alignment_trajectory(weights: Tensor, times: Tensor, durations: Tensor) -> Tensor:
|
||||||
|
"""Return expected normalized source time at each common-grid position."""
|
||||||
|
if weights.ndim != 3 or times.shape != (weights.shape[0], weights.shape[2]):
|
||||||
|
raise ValueError("weights [B,K,L] and times [B,L] must agree")
|
||||||
|
if durations.shape != (weights.shape[0],):
|
||||||
|
raise ValueError("durations must have shape [B]")
|
||||||
|
expected_seconds = torch.bmm(weights, times.unsqueeze(-1)).squeeze(-1)
|
||||||
|
return expected_seconds / durations[:, None].clamp_min(torch.finfo(expected_seconds.dtype).eps)
|
||||||
|
|
||||||
|
|
||||||
|
def monotonicity_violation_rate(trajectory: Tensor, epsilon: float = 0.02) -> Tensor:
|
||||||
|
"""Per-sample fraction of adjacent grid pairs that move backwards by epsilon."""
|
||||||
|
if trajectory.ndim != 2 or trajectory.shape[1] < 2:
|
||||||
|
raise ValueError("trajectory must have shape [B, K] with K >= 2")
|
||||||
|
if epsilon < 0:
|
||||||
|
raise ValueError("epsilon must be non-negative")
|
||||||
|
return ((trajectory[:, :-1] - trajectory[:, 1:]) > epsilon).float().mean(dim=1)
|
||||||
|
|
||||||
|
|
||||||
|
def normalized_attention_entropy(weights: Tensor, valid: Tensor) -> Tensor:
|
||||||
|
"""Per-row entropy normalized by the number of valid source positions."""
|
||||||
|
if weights.ndim != 3 or valid.shape != (weights.shape[0], weights.shape[2]):
|
||||||
|
raise ValueError("weights [B,K,L] and valid [B,L] must agree")
|
||||||
|
safe = weights.clamp_min(torch.finfo(weights.dtype).tiny)
|
||||||
|
entropy = -(weights * safe.log()).sum(dim=-1)
|
||||||
|
counts = valid.sum(dim=-1).clamp_min(1)
|
||||||
|
denominator = counts.float().log().clamp_min(torch.finfo(torch.float32).eps)
|
||||||
|
normalized = entropy / denominator[:, None]
|
||||||
|
return torch.where(counts[:, None] > 1, normalized, torch.zeros_like(normalized))
|
||||||
|
|
||||||
|
|
||||||
|
def attention_row_similarity(weights: Tensor, min_separation: int = 1) -> Tensor:
|
||||||
|
"""Mean cosine similarity between attention rows, per sample.
|
||||||
|
|
||||||
|
``min_separation=1`` compares every distinct pair (the C_row collapse
|
||||||
|
score). A value of 6 compares only pairs more than five slots apart,
|
||||||
|
matching the training diversity loss. Scores near one mean that slots
|
||||||
|
attend to nearly the same source distribution.
|
||||||
|
"""
|
||||||
|
if weights.ndim != 3:
|
||||||
|
raise ValueError("weights must have shape [B, K, L]")
|
||||||
|
if min_separation < 1:
|
||||||
|
raise ValueError("min_separation must be at least one")
|
||||||
|
batch_size, grid_size, _ = weights.shape
|
||||||
|
if grid_size <= min_separation:
|
||||||
|
return torch.zeros(batch_size, dtype=weights.dtype, device=weights.device)
|
||||||
|
normalized = F.normalize(weights, p=2, dim=-1, eps=1e-12)
|
||||||
|
similarities = torch.bmm(normalized, normalized.transpose(1, 2))
|
||||||
|
pair_mask = torch.triu(
|
||||||
|
torch.ones(grid_size, grid_size, dtype=torch.bool, device=weights.device),
|
||||||
|
diagonal=min_separation,
|
||||||
|
)
|
||||||
|
return similarities[:, pair_mask].mean(dim=-1)
|
||||||
|
|
||||||
|
|
||||||
|
def attention_width80(weights: Tensor, threshold: float = 0.8) -> Tensor:
|
||||||
|
"""Shortest contiguous source-index span containing the requested mass.
|
||||||
|
|
||||||
|
Returns integer widths with shape ``[B, K]``. This is a diagnostic, not a
|
||||||
|
score to maximize or minimize on its own.
|
||||||
|
"""
|
||||||
|
if weights.ndim != 3 or not 0 < threshold <= 1:
|
||||||
|
raise ValueError("weights must be [B,K,L] and threshold in (0, 1]")
|
||||||
|
rows = weights.detach().to(device="cpu", dtype=torch.float64).numpy()
|
||||||
|
widths = np.empty(rows.shape[:2], dtype=np.int64)
|
||||||
|
for batch_index in range(rows.shape[0]):
|
||||||
|
for grid_index in range(rows.shape[1]):
|
||||||
|
row = rows[batch_index, grid_index]
|
||||||
|
left = 0
|
||||||
|
mass = 0.0
|
||||||
|
best = len(row)
|
||||||
|
for right, value in enumerate(row):
|
||||||
|
mass += float(value)
|
||||||
|
while left <= right and mass - float(row[left]) >= threshold:
|
||||||
|
mass -= float(row[left])
|
||||||
|
left += 1
|
||||||
|
if mass + 1e-12 >= threshold:
|
||||||
|
best = min(best, right - left + 1)
|
||||||
|
widths[batch_index, grid_index] = best
|
||||||
|
return torch.from_numpy(widths)
|
||||||
|
|
||||||
|
|
||||||
|
def retrieval_metrics(query: Tensor, target: Tensor, chunk_size: int = 256) -> dict[str, float]:
|
||||||
|
"""Grid-slot retrieval; same flattened sample/slot index is the positive.
|
||||||
|
|
||||||
|
Use only as a representation-consistency probe. It is not independent
|
||||||
|
temporal ground truth; report human-labeled temporal scores separately.
|
||||||
|
"""
|
||||||
|
if query.ndim != 3 or target.ndim != 3 or query.shape != target.shape:
|
||||||
|
raise ValueError("query and target must have matching [N, K, D] shapes")
|
||||||
|
if query.shape[0] * query.shape[1] < 1:
|
||||||
|
raise ValueError("retrieval requires at least one grid position")
|
||||||
|
q = F.normalize(query.flatten(0, 1), dim=-1)
|
||||||
|
t = F.normalize(target.flatten(0, 1), dim=-1)
|
||||||
|
total = q.shape[0]
|
||||||
|
ranks = torch.empty(total, dtype=torch.long, device=q.device)
|
||||||
|
for start in range(0, total, chunk_size):
|
||||||
|
stop = min(start + chunk_size, total)
|
||||||
|
scores = q[start:stop] @ t.T
|
||||||
|
positives = scores[torch.arange(stop - start, device=q.device),
|
||||||
|
torch.arange(start, stop, device=q.device)]
|
||||||
|
ranks[start:stop] = 1 + (scores > positives[:, None]).sum(dim=1)
|
||||||
|
ranks_f = ranks.float()
|
||||||
|
return {
|
||||||
|
"r_at_1": float((ranks <= 1).float().mean().item()),
|
||||||
|
"r_at_5": float((ranks <= min(5, total)).float().mean().item()),
|
||||||
|
"mrr": float((1.0 / ranks_f).mean().item()),
|
||||||
|
"queries": float(total),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def summarize_alignment(
|
||||||
|
weights: dict[str, Tensor],
|
||||||
|
times: dict[str, Tensor],
|
||||||
|
valid: dict[str, Tensor],
|
||||||
|
durations: Tensor,
|
||||||
|
epsilon: float = 0.02,
|
||||||
|
) -> dict[str, dict[str, float]]:
|
||||||
|
"""Produce sample-aggregated E1-E3 diagnostics for each modality."""
|
||||||
|
summary: dict[str, dict[str, float]] = {}
|
||||||
|
for name in MODALITIES:
|
||||||
|
trajectory = alignment_trajectory(weights[name], times[name], durations)
|
||||||
|
mvr = monotonicity_violation_rate(trajectory, epsilon)
|
||||||
|
entropy = normalized_attention_entropy(weights[name], valid[name])
|
||||||
|
width = attention_width80(weights[name])
|
||||||
|
summary[name] = {
|
||||||
|
"mvr": float(mvr.mean().item()),
|
||||||
|
"normalized_entropy": float(entropy.mean().item()),
|
||||||
|
"width80_indices": float(width.float().mean().item()),
|
||||||
|
"mean_time_start": float(trajectory[:, 0].mean().item()),
|
||||||
|
"mean_time_end": float(trajectory[:, -1].mean().item()),
|
||||||
|
"trajectory_span_fraction": float(
|
||||||
|
(trajectory[:, -1] - trajectory[:, 0]).mean().item()
|
||||||
|
),
|
||||||
|
}
|
||||||
|
return summary
|
||||||
@@ -0,0 +1,247 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from collections.abc import Mapping
|
||||||
|
|
||||||
|
import torch
|
||||||
|
from torch import Tensor, nn
|
||||||
|
|
||||||
|
from .alignment import index_alignment
|
||||||
|
from .types import AlignmentOutput, MODALITIES, SequenceBatch
|
||||||
|
|
||||||
|
|
||||||
|
def _sinusoidal_position_encoding(length: int, dimension: int) -> Tensor:
|
||||||
|
"""Build a fixed, small-amplitude sinusoidal code for ordered slots."""
|
||||||
|
positions = torch.arange(length, dtype=torch.float32).unsqueeze(1)
|
||||||
|
frequencies = torch.exp(
|
||||||
|
torch.arange(0, dimension, 2, dtype=torch.float32)
|
||||||
|
* (-torch.log(torch.tensor(10000.0)) / dimension)
|
||||||
|
)
|
||||||
|
encoding = torch.zeros(length, dimension, dtype=torch.float32)
|
||||||
|
encoding[:, 0::2] = torch.sin(positions * frequencies)
|
||||||
|
odd_width = encoding[:, 1::2].shape[1]
|
||||||
|
if odd_width:
|
||||||
|
encoding[:, 1::2] = torch.cos(positions * frequencies[:odd_width])
|
||||||
|
return encoding * (dimension**-0.5)
|
||||||
|
|
||||||
|
|
||||||
|
def _temporal_position_encoding(times: Tensor, dimension: int) -> Tensor:
|
||||||
|
"""Encode normalized source times with a fixed Fourier feature bank."""
|
||||||
|
half_width = (dimension + 1) // 2
|
||||||
|
frequencies = torch.logspace(
|
||||||
|
0.0,
|
||||||
|
1.6989700043360187,
|
||||||
|
steps=half_width,
|
||||||
|
device=times.device,
|
||||||
|
dtype=times.dtype,
|
||||||
|
)
|
||||||
|
angles = (2.0 * torch.pi) * times.unsqueeze(-1) * frequencies
|
||||||
|
encoding = torch.empty(*times.shape, dimension, device=times.device, dtype=times.dtype)
|
||||||
|
encoding[..., 0::2] = torch.sin(angles)
|
||||||
|
if dimension > 1:
|
||||||
|
encoding[..., 1::2] = torch.cos(angles[..., : encoding[..., 1::2].shape[-1]])
|
||||||
|
return encoding * (dimension**-0.25)
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_inputs(sequences: Mapping[str, SequenceBatch], grid_size: int) -> int:
|
||||||
|
if set(sequences) != set(MODALITIES):
|
||||||
|
raise ValueError(f"sequences must contain exactly {MODALITIES}")
|
||||||
|
batch_sizes = {sequences[name].features.shape[0] for name in MODALITIES}
|
||||||
|
if len(batch_sizes) != 1:
|
||||||
|
raise ValueError("all modalities must have the same batch size")
|
||||||
|
if grid_size < 1:
|
||||||
|
raise ValueError("grid_size must be positive")
|
||||||
|
return batch_sizes.pop()
|
||||||
|
|
||||||
|
|
||||||
|
class _CrossAttention(nn.Module):
|
||||||
|
def __init__(self, dimension: int, heads: int, dropout: float) -> None:
|
||||||
|
super().__init__()
|
||||||
|
if dimension % heads != 0:
|
||||||
|
raise ValueError("dimension must be divisible by heads")
|
||||||
|
self.attention = nn.MultiheadAttention(
|
||||||
|
# Keep the returned alignment matrix row-stochastic during training.
|
||||||
|
# PyTorch applies attention dropout to returned weights when it is
|
||||||
|
# nonzero, which breaks the shared AlignmentOutput contract.
|
||||||
|
embed_dim=dimension, num_heads=heads, dropout=0.0, batch_first=True
|
||||||
|
)
|
||||||
|
self.input_dropout = nn.Dropout(dropout)
|
||||||
|
|
||||||
|
def forward(
|
||||||
|
self,
|
||||||
|
query: Tensor,
|
||||||
|
source: Tensor,
|
||||||
|
valid: Tensor,
|
||||||
|
*,
|
||||||
|
source_position: Tensor | None = None,
|
||||||
|
) -> tuple[Tensor, Tensor]:
|
||||||
|
key = source if source_position is None else source + source_position
|
||||||
|
values, weights = self.attention(
|
||||||
|
self.input_dropout(query),
|
||||||
|
self.input_dropout(key),
|
||||||
|
self.input_dropout(source),
|
||||||
|
key_padding_mask=~valid,
|
||||||
|
need_weights=True,
|
||||||
|
average_attn_weights=True,
|
||||||
|
)
|
||||||
|
return values, weights
|
||||||
|
|
||||||
|
|
||||||
|
class TextAnchoredCrossAttention(nn.Module):
|
||||||
|
"""M3: transcript-order text slots query Audio and Vision sequences."""
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
dimensions: Mapping[str, int],
|
||||||
|
grid_size: int = 50,
|
||||||
|
hidden_size: int = 128,
|
||||||
|
heads: int = 4,
|
||||||
|
dropout: float = 0.1,
|
||||||
|
source_time_encoding: bool = False,
|
||||||
|
) -> None:
|
||||||
|
super().__init__()
|
||||||
|
if set(dimensions) != set(MODALITIES):
|
||||||
|
raise ValueError(f"dimensions must contain exactly {MODALITIES}")
|
||||||
|
self.grid_size = grid_size
|
||||||
|
self.source_time_encoding = source_time_encoding
|
||||||
|
self.hidden_size = hidden_size
|
||||||
|
self.projections = nn.ModuleDict(
|
||||||
|
{name: nn.Linear(dimensions[name], hidden_size) for name in MODALITIES}
|
||||||
|
)
|
||||||
|
self.audio_attention = _CrossAttention(hidden_size, heads, dropout)
|
||||||
|
self.vision_attention = _CrossAttention(hidden_size, heads, dropout)
|
||||||
|
|
||||||
|
def forward(
|
||||||
|
self, sequences: Mapping[str, SequenceBatch], durations: Tensor | None = None
|
||||||
|
) -> AlignmentOutput:
|
||||||
|
_validate_inputs(sequences, self.grid_size)
|
||||||
|
if self.source_time_encoding:
|
||||||
|
if durations is None or durations.shape != (sequences["text"].features.shape[0],):
|
||||||
|
raise ValueError("durations with shape [B] are required for source time encoding")
|
||||||
|
durations = durations.to(device=sequences["text"].times.device).clamp_min(1e-8)
|
||||||
|
projected = {
|
||||||
|
name: self.projections[name](sequences[name].features) for name in MODALITIES
|
||||||
|
}
|
||||||
|
text_weights, text_fallbacks = index_alignment(sequences["text"].valid, self.grid_size)
|
||||||
|
text_query = torch.bmm(text_weights.to(projected["text"].dtype), projected["text"])
|
||||||
|
|
||||||
|
source_positions: dict[str, Tensor] = {}
|
||||||
|
if self.source_time_encoding:
|
||||||
|
assert durations is not None
|
||||||
|
text_centers = torch.bmm(
|
||||||
|
text_weights.to(sequences["text"].times.dtype),
|
||||||
|
sequences["text"].times.unsqueeze(-1),
|
||||||
|
).squeeze(-1) / durations[:, None]
|
||||||
|
text_query = text_query + _temporal_position_encoding(
|
||||||
|
text_centers, self.hidden_size
|
||||||
|
).to(text_query.dtype)
|
||||||
|
for name in ("audio", "vision"):
|
||||||
|
normalized_times = sequences[name].times / durations[:, None]
|
||||||
|
source_positions[name] = _temporal_position_encoding(
|
||||||
|
normalized_times, self.hidden_size
|
||||||
|
).to(projected[name].dtype)
|
||||||
|
|
||||||
|
audio_values, audio_weights = self.audio_attention(
|
||||||
|
text_query,
|
||||||
|
projected["audio"],
|
||||||
|
sequences["audio"].valid,
|
||||||
|
source_position=source_positions.get("audio"),
|
||||||
|
)
|
||||||
|
vision_values, vision_weights = self.vision_attention(
|
||||||
|
text_query,
|
||||||
|
projected["vision"],
|
||||||
|
sequences["vision"].valid,
|
||||||
|
source_position=source_positions.get("vision"),
|
||||||
|
)
|
||||||
|
output = AlignmentOutput(
|
||||||
|
weights={"text": text_weights, "audio": audio_weights, "vision": vision_weights},
|
||||||
|
aligned={"text": text_query, "audio": audio_values, "vision": vision_values},
|
||||||
|
fallback_rows={"text": text_fallbacks, "audio": 0, "vision": 0},
|
||||||
|
)
|
||||||
|
output.validate({name: sequences[name].valid for name in MODALITIES})
|
||||||
|
return output
|
||||||
|
|
||||||
|
|
||||||
|
class SharedLatentTimeline(nn.Module):
|
||||||
|
"""M4: K learned shared slots attend independently to all three modalities."""
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
dimensions: Mapping[str, int],
|
||||||
|
grid_size: int = 50,
|
||||||
|
hidden_size: int = 128,
|
||||||
|
heads: int = 4,
|
||||||
|
dropout: float = 0.1,
|
||||||
|
absolute_position_encoding: bool = False,
|
||||||
|
source_time_encoding: bool = False,
|
||||||
|
) -> None:
|
||||||
|
super().__init__()
|
||||||
|
if set(dimensions) != set(MODALITIES):
|
||||||
|
raise ValueError(f"dimensions must contain exactly {MODALITIES}")
|
||||||
|
if hidden_size % heads != 0:
|
||||||
|
raise ValueError("hidden_size must be divisible by heads")
|
||||||
|
self.grid_size = grid_size
|
||||||
|
self.hidden_size = hidden_size
|
||||||
|
self.absolute_position_encoding = absolute_position_encoding
|
||||||
|
self.source_time_encoding = source_time_encoding
|
||||||
|
self.register_buffer(
|
||||||
|
"sinusoidal_positions",
|
||||||
|
_sinusoidal_position_encoding(grid_size, hidden_size),
|
||||||
|
persistent=False,
|
||||||
|
)
|
||||||
|
self.projections = nn.ModuleDict(
|
||||||
|
{name: nn.Linear(dimensions[name], hidden_size) for name in MODALITIES}
|
||||||
|
)
|
||||||
|
self.slots = nn.Parameter(torch.empty(grid_size, hidden_size))
|
||||||
|
nn.init.normal_(self.slots, mean=0.0, std=hidden_size**-0.5)
|
||||||
|
self.attention = nn.ModuleDict(
|
||||||
|
{name: _CrossAttention(hidden_size, heads, dropout) for name in MODALITIES}
|
||||||
|
)
|
||||||
|
|
||||||
|
def forward(
|
||||||
|
self, sequences: Mapping[str, SequenceBatch], durations: Tensor | None = None
|
||||||
|
) -> AlignmentOutput:
|
||||||
|
batch_size = _validate_inputs(sequences, self.grid_size)
|
||||||
|
if self.source_time_encoding:
|
||||||
|
if durations is None or durations.shape != (batch_size,):
|
||||||
|
raise ValueError("durations with shape [B] are required for source time encoding")
|
||||||
|
durations = durations.to(device=sequences["text"].times.device).clamp_min(1e-8)
|
||||||
|
latent_queries = self.slots.unsqueeze(0).expand(batch_size, -1, -1)
|
||||||
|
if self.absolute_position_encoding:
|
||||||
|
latent_queries = latent_queries + self.sinusoidal_positions.unsqueeze(0)
|
||||||
|
if self.source_time_encoding:
|
||||||
|
centers = (
|
||||||
|
torch.arange(
|
||||||
|
self.grid_size,
|
||||||
|
dtype=sequences["text"].times.dtype,
|
||||||
|
device=sequences["text"].times.device,
|
||||||
|
)
|
||||||
|
+ 0.5
|
||||||
|
) / self.grid_size
|
||||||
|
centers = centers.unsqueeze(0).expand(batch_size, -1)
|
||||||
|
latent_queries = latent_queries + _temporal_position_encoding(
|
||||||
|
centers, self.hidden_size
|
||||||
|
).to(latent_queries.dtype)
|
||||||
|
weights: dict[str, Tensor] = {}
|
||||||
|
aligned: dict[str, Tensor] = {}
|
||||||
|
for name in MODALITIES:
|
||||||
|
source = self.projections[name](sequences[name].features)
|
||||||
|
source_position = None
|
||||||
|
if self.source_time_encoding:
|
||||||
|
assert durations is not None
|
||||||
|
normalized_times = sequences[name].times / durations[:, None]
|
||||||
|
source_position = _temporal_position_encoding(
|
||||||
|
normalized_times, self.hidden_size
|
||||||
|
).to(source.dtype)
|
||||||
|
aligned[name], weights[name] = self.attention[name](
|
||||||
|
latent_queries,
|
||||||
|
source,
|
||||||
|
sequences[name].valid,
|
||||||
|
source_position=source_position,
|
||||||
|
)
|
||||||
|
output = AlignmentOutput(
|
||||||
|
weights=weights,
|
||||||
|
aligned=aligned,
|
||||||
|
fallback_rows={name: 0 for name in MODALITIES},
|
||||||
|
)
|
||||||
|
output.validate({name: sequences[name].valid for name in MODALITIES})
|
||||||
|
return output
|
||||||
@@ -0,0 +1,194 @@
|
|||||||
|
"""Summarize the saved D0-D3 alignment-diagnostic metrics without retraining."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import csv
|
||||||
|
from collections import defaultdict
|
||||||
|
from pathlib import Path
|
||||||
|
import shutil
|
||||||
|
from statistics import mean
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
METRICS = (
|
||||||
|
"mvr",
|
||||||
|
"normalized_entropy",
|
||||||
|
"c_row",
|
||||||
|
"trajectory_span",
|
||||||
|
"mean_absolute_time_center_error",
|
||||||
|
"gaussian_target_kl",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def merge_trial_metrics(run_dir: Path, destination: Path) -> None:
|
||||||
|
"""Rebuild the combined file with a variant label for PE controls."""
|
||||||
|
rows: list[dict[str, str]] = []
|
||||||
|
fields: list[str] | None = None
|
||||||
|
for experiment in ("D0", "D1", "D2", "D3"):
|
||||||
|
for path in sorted((run_dir / experiment).glob("*_metrics.csv")):
|
||||||
|
variant = path.name.removesuffix("_metrics.csv")
|
||||||
|
with path.open(newline="", encoding="utf-8-sig") as handle:
|
||||||
|
reader = csv.DictReader(handle)
|
||||||
|
if fields is None:
|
||||||
|
fields = ["experiment", "method", "variant"] + [
|
||||||
|
name for name in (reader.fieldnames or [])
|
||||||
|
if name not in {"experiment", "method", "variant"}
|
||||||
|
]
|
||||||
|
for row in reader:
|
||||||
|
row["variant"] = variant
|
||||||
|
rows.append(row)
|
||||||
|
if not fields:
|
||||||
|
raise FileNotFoundError(f"no per-trial metric CSV files found under {run_dir}")
|
||||||
|
destination.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
with destination.open("w", newline="", encoding="utf-8-sig") as handle:
|
||||||
|
writer = csv.DictWriter(handle, fieldnames=fields)
|
||||||
|
writer.writeheader()
|
||||||
|
writer.writerows(rows)
|
||||||
|
|
||||||
|
|
||||||
|
def summarize(source: Path, destination: Path) -> list[dict[str, Any]]:
|
||||||
|
groups: dict[tuple[str, str, str], list[dict[str, str]]] = defaultdict(list)
|
||||||
|
with source.open(newline="", encoding="utf-8-sig") as handle:
|
||||||
|
for row in csv.DictReader(handle):
|
||||||
|
if row["experiment"] in {"D2", "D3"}:
|
||||||
|
groups[(row["experiment"], row["method"], row["modality"])].append(row)
|
||||||
|
|
||||||
|
summaries: list[dict[str, Any]] = []
|
||||||
|
for (experiment, method, modality), rows in sorted(groups.items()):
|
||||||
|
summary: dict[str, Any] = {
|
||||||
|
"experiment": experiment,
|
||||||
|
"method": method,
|
||||||
|
"modality": modality,
|
||||||
|
"sample_count": len(rows),
|
||||||
|
}
|
||||||
|
for metric in METRICS:
|
||||||
|
values = [float(row[metric]) for row in rows if row.get(metric, "") != ""]
|
||||||
|
summary[f"mean_{metric}"] = mean(values) if values else ""
|
||||||
|
summary[f"n_{metric}"] = len(values)
|
||||||
|
summaries.append(summary)
|
||||||
|
|
||||||
|
destination.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
with destination.open("w", newline="", encoding="utf-8-sig") as handle:
|
||||||
|
writer = csv.DictWriter(handle, fieldnames=list(summaries[0]))
|
||||||
|
writer.writeheader()
|
||||||
|
writer.writerows(summaries)
|
||||||
|
return summaries
|
||||||
|
|
||||||
|
|
||||||
|
def build_report_bundle(run_dir: Path) -> Path:
|
||||||
|
"""Copy compact diagnostic evidence, leaving large checkpoints local."""
|
||||||
|
bundle = run_dir / "report_bundle"
|
||||||
|
bundle.mkdir(parents=True, exist_ok=True)
|
||||||
|
for name in (
|
||||||
|
"debug_summary.csv",
|
||||||
|
"metric_summary.csv",
|
||||||
|
"per_sample_metrics.csv",
|
||||||
|
"run_manifest.json",
|
||||||
|
"experiment.log",
|
||||||
|
):
|
||||||
|
source = run_dir / name
|
||||||
|
if source.exists():
|
||||||
|
shutil.copy2(source, bundle / name)
|
||||||
|
for experiment in ("D0", "D1", "D2", "D3"):
|
||||||
|
source_dir = run_dir / experiment
|
||||||
|
target_dir = bundle / experiment
|
||||||
|
target_dir.mkdir(exist_ok=True)
|
||||||
|
for pattern in ("*_history.csv", "*_metrics.csv", "*_heatmap.png", "*_trajectory.png"):
|
||||||
|
for source in source_dir.glob(pattern):
|
||||||
|
shutil.copy2(source, target_dir / source.name)
|
||||||
|
synthetic_source = run_dir / "synthetic"
|
||||||
|
synthetic_target = bundle / "synthetic"
|
||||||
|
synthetic_target.mkdir(exist_ok=True)
|
||||||
|
for name in ("metrics.csv", "training_history.csv", "run_manifest.json"):
|
||||||
|
source = synthetic_source / name
|
||||||
|
if source.exists():
|
||||||
|
shutil.copy2(source, synthetic_target / name)
|
||||||
|
for source in synthetic_source.rglob("*.png"):
|
||||||
|
target = synthetic_target / source.relative_to(synthetic_source)
|
||||||
|
target.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
shutil.copy2(source, target)
|
||||||
|
source_time_dir = run_dir / "source_time"
|
||||||
|
source_time_target = bundle / "source_time"
|
||||||
|
source_time_target.mkdir(exist_ok=True)
|
||||||
|
for name in ("summary.csv", "per_sample_metrics.csv", "run_manifest.json"):
|
||||||
|
source = source_time_dir / name
|
||||||
|
if source.exists():
|
||||||
|
shutil.copy2(source, source_time_target / name)
|
||||||
|
for variant_dir in source_time_dir.glob("M*"):
|
||||||
|
target_dir = source_time_target / variant_dir.name
|
||||||
|
target_dir.mkdir(exist_ok=True)
|
||||||
|
for pattern in ("history.csv", "metrics.csv", "*_heatmap.png", "*_trajectory.png"):
|
||||||
|
for source in variant_dir.glob(pattern):
|
||||||
|
shutil.copy2(source, target_dir / source.name)
|
||||||
|
heldout_dir = run_dir / "heldout"
|
||||||
|
heldout_target = bundle / "heldout"
|
||||||
|
heldout_target.mkdir(exist_ok=True)
|
||||||
|
for name in (
|
||||||
|
"heldout_summary.csv",
|
||||||
|
"per_sample_metrics.csv",
|
||||||
|
"training_summary.csv",
|
||||||
|
"training_history.csv",
|
||||||
|
"run_manifest.json",
|
||||||
|
):
|
||||||
|
source = heldout_dir / name
|
||||||
|
if source.exists():
|
||||||
|
shutil.copy2(source, heldout_target / name)
|
||||||
|
for variant_dir in heldout_dir.glob("M*"):
|
||||||
|
target_dir = heldout_target / variant_dir.name
|
||||||
|
target_dir.mkdir(exist_ok=True)
|
||||||
|
for pattern in ("history.csv", "heldout_metrics.csv", "*_heatmap.png", "*_trajectory.png"):
|
||||||
|
for source in variant_dir.glob(pattern):
|
||||||
|
shutil.copy2(source, target_dir / source.name)
|
||||||
|
readme = """# Q1 M3/M4 可学习性诊断结果包
|
||||||
|
|
||||||
|
本包包含 D0–D3 诊断的汇总表、运行清单、日志、各试验的损失历史、逐样本指标和代表性图像。PyTorch 检查点与逐样本 `.npz` 注意力矩阵留在上级 `outputs/alignment_debug/`,因此结果包较轻。
|
||||||
|
|
||||||
|
- D0/D1:单样本过拟合诊断。
|
||||||
|
- D1-S:输入仅为合成时间坐标;M3/M4 都能把 Gaussian 目标拟合至约 1e-4 KL。
|
||||||
|
- D2/D3:100 条样本上的同集训练/评价诊断,单个随机种子,不用于声称泛化。
|
||||||
|
- D4:真实单样本加入显式时间 key/query 特征后,Audio 注意力形成局部时间带。
|
||||||
|
- D5:按 `video_id` 留出 20 条样本;M4 的 Audio/Vision 时间带能迁移,M3 有改善但仍未充分贴近目标。
|
||||||
|
- Gaussian 时间目标是根据源时间戳生成的弱先验,不是人工对齐真值。
|
||||||
|
- M3 的 Gaussian KL 取 Audio/Vision 平均,M4 取 Text/Audio/Vision 平均;KL 数值仅用于各自的优化诊断,不可作为方法排名。
|
||||||
|
- D1 真实特征上 Text/Vision 能拟合,Audio 仍失败;合成对照成功,提示真实源特征的显式时间身份值得优先验证。
|
||||||
|
- 全数据 Q/K 梯度仍非零,但学习到的注意力没有稳定形成局部时间带;加入重构和对比目标后 Audio 塌缩更明显。
|
||||||
|
- 环境:Fedora WSL,`uv`,NVIDIA GeForce RTX 5070 Ti,Python 3.14.7,PyTorch 2.14.0+cu130。
|
||||||
|
|
||||||
|
详细解释见项目根目录的 `RESULTS.md`。
|
||||||
|
"""
|
||||||
|
(bundle / "README.md").write_text(readme, encoding="utf-8")
|
||||||
|
return bundle
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
project = Path(__file__).resolve().parents[1]
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument(
|
||||||
|
"--input",
|
||||||
|
type=Path,
|
||||||
|
default=project / "outputs/alignment_debug/per_sample_metrics.csv",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--output",
|
||||||
|
type=Path,
|
||||||
|
default=project / "outputs/alignment_debug/metric_summary.csv",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
merge_trial_metrics(args.input.parent, args.input)
|
||||||
|
for row in summarize(args.input, args.output):
|
||||||
|
values = " ".join(
|
||||||
|
f"{name}={row[f'mean_{name}']:.4f}"
|
||||||
|
for name in METRICS
|
||||||
|
if row[f"mean_{name}"] != ""
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f"{row['experiment']} {row['method']} {row['modality']} "
|
||||||
|
f"n={row['sample_count']} {values}"
|
||||||
|
)
|
||||||
|
print(f"Wrote {args.output}")
|
||||||
|
print(f"Built report bundle at {build_report_bundle(args.input.parent)}")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,338 @@
|
|||||||
|
"""Fit Gaussian time bands with synthetic time-only inputs as an attention control."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
import platform
|
||||||
|
import random
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import matplotlib
|
||||||
|
|
||||||
|
matplotlib.use("Agg")
|
||||||
|
import matplotlib.pyplot as plt
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
from torch import Tensor, nn
|
||||||
|
|
||||||
|
from .metrics import alignment_trajectory, normalized_attention_entropy
|
||||||
|
from .models import SharedLatentTimeline, TextAnchoredCrossAttention
|
||||||
|
from .types import MODALITIES, SequenceBatch
|
||||||
|
|
||||||
|
|
||||||
|
GRID_SIZE = 50
|
||||||
|
HIDDEN_SIZE = 128
|
||||||
|
SIGMA = 0.10
|
||||||
|
SEED = 42
|
||||||
|
MODALITY_LENGTHS = {"text": GRID_SIZE, "audio": 256, "vision": 100}
|
||||||
|
|
||||||
|
|
||||||
|
def _seed(seed: int) -> None:
|
||||||
|
random.seed(seed)
|
||||||
|
np.random.seed(seed)
|
||||||
|
torch.manual_seed(seed)
|
||||||
|
if torch.cuda.is_available():
|
||||||
|
torch.cuda.manual_seed_all(seed)
|
||||||
|
torch.backends.cudnn.deterministic = True
|
||||||
|
torch.backends.cudnn.benchmark = False
|
||||||
|
|
||||||
|
|
||||||
|
def _synthetic_sequence(length: int, device: torch.device) -> SequenceBatch:
|
||||||
|
times = torch.linspace(0.0, 1.0, length, device=device).unsqueeze(0)
|
||||||
|
# These features contain only the source's synthetic time coordinate.
|
||||||
|
features = torch.stack(
|
||||||
|
(times, times.square(), torch.ones_like(times)), dim=-1
|
||||||
|
)
|
||||||
|
valid = torch.ones_like(times, dtype=torch.bool)
|
||||||
|
return SequenceBatch(features=features, times=times, valid=valid)
|
||||||
|
|
||||||
|
|
||||||
|
def _targets(
|
||||||
|
sequences: dict[str, SequenceBatch], method: str, device: torch.device
|
||||||
|
) -> dict[str, Tensor]:
|
||||||
|
centers = (torch.arange(GRID_SIZE, device=device, dtype=torch.float32) + 0.5) / GRID_SIZE
|
||||||
|
names = ("audio", "vision") if method == "M3" else MODALITIES
|
||||||
|
targets: dict[str, Tensor] = {}
|
||||||
|
for name in names:
|
||||||
|
times = sequences[name].times
|
||||||
|
logits = -0.5 * ((times[:, None, :] - centers[None, :, None]) / SIGMA).square()
|
||||||
|
targets[name] = torch.softmax(logits, dim=-1)
|
||||||
|
return targets
|
||||||
|
|
||||||
|
|
||||||
|
def _loss(output: Any, targets: dict[str, Tensor]) -> Tensor:
|
||||||
|
losses = []
|
||||||
|
for name, target in targets.items():
|
||||||
|
predicted = output.weights[name].clamp_min(1e-8)
|
||||||
|
safe_target = target.clamp_min(1e-12)
|
||||||
|
losses.append((safe_target * (safe_target.log() - predicted.log())).sum(-1).mean())
|
||||||
|
return torch.stack(losses).mean()
|
||||||
|
|
||||||
|
|
||||||
|
def _attention_layers(model: nn.Module, method: str) -> list[nn.MultiheadAttention]:
|
||||||
|
if method == "M3":
|
||||||
|
return [model.audio_attention.attention, model.vision_attention.attention]
|
||||||
|
return [model.attention[name].attention for name in MODALITIES]
|
||||||
|
|
||||||
|
|
||||||
|
def _gradient_summary(
|
||||||
|
model: nn.Module, method: str, loss: Tensor
|
||||||
|
) -> dict[str, float]:
|
||||||
|
layers = _attention_layers(model, method)
|
||||||
|
params = [layer.in_proj_weight for layer in layers]
|
||||||
|
if method == "M4":
|
||||||
|
params.append(model.slots)
|
||||||
|
gradients = torch.autograd.grad(loss, params, retain_graph=True)
|
||||||
|
q_sq = torch.zeros((), device=loss.device)
|
||||||
|
k_sq = torch.zeros((), device=loss.device)
|
||||||
|
for grad in gradients[: len(layers)]:
|
||||||
|
q_sq += grad[:HIDDEN_SIZE].square().sum()
|
||||||
|
k_sq += grad[HIDDEN_SIZE : 2 * HIDDEN_SIZE].square().sum()
|
||||||
|
result = {"grad_WQ": float(q_sq.sqrt().item()), "grad_WK": float(k_sq.sqrt().item())}
|
||||||
|
if method == "M4":
|
||||||
|
result["grad_Z"] = float(gradients[-1].norm().item())
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
def _write_csv(path: Path, rows: list[dict[str, Any]]) -> None:
|
||||||
|
if not rows:
|
||||||
|
return
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
with path.open("w", newline="", encoding="utf-8-sig") as handle:
|
||||||
|
fields = list(dict.fromkeys(key for row in rows for key in row))
|
||||||
|
writer = csv.DictWriter(handle, fieldnames=fields)
|
||||||
|
writer.writeheader()
|
||||||
|
writer.writerows(rows)
|
||||||
|
|
||||||
|
|
||||||
|
def _save_plots(
|
||||||
|
folder: Path,
|
||||||
|
method: str,
|
||||||
|
variant: str,
|
||||||
|
output: Any,
|
||||||
|
targets: dict[str, Tensor],
|
||||||
|
sequences: dict[str, SequenceBatch],
|
||||||
|
duration: float = 1.0,
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
names = ("audio", "vision") if method == "M3" else MODALITIES
|
||||||
|
rows: list[dict[str, Any]] = []
|
||||||
|
fig, axes = plt.subplots(len(names), 2, figsize=(11, 3.4 * len(names)), constrained_layout=True)
|
||||||
|
if len(names) == 1:
|
||||||
|
axes = np.asarray([axes])
|
||||||
|
fig_traj, ax_traj = plt.subplots(figsize=(8, 5), constrained_layout=True)
|
||||||
|
grid = (np.arange(GRID_SIZE, dtype=np.float32) + 0.5) / GRID_SIZE
|
||||||
|
for row_index, name in enumerate(names):
|
||||||
|
weights = output.weights[name][0].detach().cpu().numpy()
|
||||||
|
target = targets[name][0].detach().cpu().numpy()
|
||||||
|
times = sequences[name].times[0].detach().cpu().numpy()
|
||||||
|
entropy = normalized_attention_entropy(
|
||||||
|
output.weights[name], sequences[name].valid
|
||||||
|
).mean().item()
|
||||||
|
trajectory = alignment_trajectory(
|
||||||
|
output.weights[name], sequences[name].times, torch.tensor([duration], device=output.weights[name].device)
|
||||||
|
)[0]
|
||||||
|
trajectory_np = trajectory.detach().cpu().numpy()
|
||||||
|
span = float(trajectory_np[-1] - trajectory_np[0])
|
||||||
|
kl = float(
|
||||||
|
(targets[name] * (targets[name].clamp_min(1e-12).log() - output.weights[name].clamp_min(1e-8).log()))
|
||||||
|
.sum(-1)
|
||||||
|
.mean()
|
||||||
|
.item()
|
||||||
|
)
|
||||||
|
rows.append(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"variant": variant,
|
||||||
|
"modality": name,
|
||||||
|
"normalized_entropy": entropy,
|
||||||
|
"trajectory_span": span,
|
||||||
|
"mean_absolute_time_center_error": float(
|
||||||
|
np.mean(np.abs(trajectory_np - grid))
|
||||||
|
),
|
||||||
|
"gaussian_target_kl": kl,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
extent = (float(times[0]), float(times[-1]), 0.0, 1.0)
|
||||||
|
image = axes[row_index, 0].imshow(
|
||||||
|
weights, origin="lower", aspect="auto", interpolation="nearest", extent=extent, cmap="magma"
|
||||||
|
)
|
||||||
|
axes[row_index, 0].set_title(f"{name}: learned A")
|
||||||
|
axes[row_index, 0].set_xlabel("synthetic source time")
|
||||||
|
axes[row_index, 0].set_ylabel("slot / K")
|
||||||
|
fig.colorbar(image, ax=axes[row_index, 0], fraction=0.046, pad=0.04)
|
||||||
|
image_target = axes[row_index, 1].imshow(
|
||||||
|
target, origin="lower", aspect="auto", interpolation="nearest", extent=extent, cmap="magma"
|
||||||
|
)
|
||||||
|
axes[row_index, 1].set_title(f"{name}: Gaussian target P")
|
||||||
|
axes[row_index, 1].set_xlabel("synthetic source time")
|
||||||
|
axes[row_index, 1].set_ylabel("slot / K")
|
||||||
|
fig.colorbar(image_target, ax=axes[row_index, 1], fraction=0.046, pad=0.04)
|
||||||
|
ax_traj.plot(grid, trajectory_np, label=f"{name} learned")
|
||||||
|
ax_traj.plot([0, 1], [0, 1], "k:", label="uniform-time reference")
|
||||||
|
ax_traj.set(xlim=(0, 1), ylim=(0, 1), xlabel="slot position", ylabel="expected source time")
|
||||||
|
ax_traj.grid(alpha=0.2)
|
||||||
|
ax_traj.legend()
|
||||||
|
ax_traj.set_title(f"{method} {variant}: synthetic alignment trajectory")
|
||||||
|
folder.mkdir(parents=True, exist_ok=True)
|
||||||
|
fig.savefig(folder / f"{variant}_heatmap.png", dpi=160)
|
||||||
|
fig_traj.savefig(folder / f"{variant}_trajectory.png", dpi=160)
|
||||||
|
plt.close(fig)
|
||||||
|
plt.close(fig_traj)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def _run_trial(
|
||||||
|
method: str,
|
||||||
|
variant: str,
|
||||||
|
*,
|
||||||
|
device: torch.device,
|
||||||
|
steps: int,
|
||||||
|
seed: int,
|
||||||
|
output_dir: Path,
|
||||||
|
) -> tuple[list[dict[str, Any]], list[dict[str, Any]], dict[str, Any]]:
|
||||||
|
_seed(seed)
|
||||||
|
dimensions = {name: 3 for name in MODALITIES}
|
||||||
|
if method == "M3":
|
||||||
|
model: nn.Module = TextAnchoredCrossAttention(
|
||||||
|
dimensions, grid_size=GRID_SIZE, hidden_size=HIDDEN_SIZE, heads=4, dropout=0.0
|
||||||
|
)
|
||||||
|
use_pe = False
|
||||||
|
else:
|
||||||
|
use_pe = variant == "M4_sinPE"
|
||||||
|
model = SharedLatentTimeline(
|
||||||
|
dimensions,
|
||||||
|
grid_size=GRID_SIZE,
|
||||||
|
hidden_size=HIDDEN_SIZE,
|
||||||
|
heads=4,
|
||||||
|
dropout=0.0,
|
||||||
|
absolute_position_encoding=use_pe,
|
||||||
|
)
|
||||||
|
model.to(device)
|
||||||
|
sequences = {
|
||||||
|
name: _synthetic_sequence(MODALITY_LENGTHS[name], device) for name in MODALITIES
|
||||||
|
}
|
||||||
|
targets = _targets(sequences, method, device)
|
||||||
|
optimizer = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=0.0)
|
||||||
|
history: list[dict[str, Any]] = []
|
||||||
|
output = None
|
||||||
|
print(f"[{variant}] synthetic time-only inputs, steps={steps}, device={device}", flush=True)
|
||||||
|
for step in range(1, steps + 1):
|
||||||
|
model.train()
|
||||||
|
output = model(sequences)
|
||||||
|
loss = _loss(output, targets)
|
||||||
|
if not torch.isfinite(loss):
|
||||||
|
raise FloatingPointError(f"non-finite synthetic alignment loss for {variant} at step {step}")
|
||||||
|
row: dict[str, Any] = {"method": method, "variant": variant, "step": step, "L_align": float(loss.item())}
|
||||||
|
if step == 1 or step % 100 == 0 or step == steps:
|
||||||
|
row.update(_gradient_summary(model, method, loss))
|
||||||
|
optimizer.zero_grad(set_to_none=True)
|
||||||
|
loss.backward()
|
||||||
|
nn.utils.clip_grad_norm_(model.parameters(), 2.0)
|
||||||
|
optimizer.step()
|
||||||
|
history.append(row)
|
||||||
|
if step == 1 or step % 100 == 0 or step == steps:
|
||||||
|
print(
|
||||||
|
f"[{variant} {step}/{steps}] KL={row['L_align']:.5f} "
|
||||||
|
f"grad_Q/K/Z={row.get('grad_WQ', 0):.3g}/{row.get('grad_WK', 0):.3g}/"
|
||||||
|
f"{row.get('grad_Z', float('nan')):.3g}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
assert output is not None
|
||||||
|
model.eval()
|
||||||
|
with torch.no_grad():
|
||||||
|
output = model(sequences)
|
||||||
|
metric_rows = _save_plots(
|
||||||
|
output_dir / variant, method, variant, output, targets, sequences
|
||||||
|
)
|
||||||
|
history_rows = history
|
||||||
|
checkpoint = output_dir / f"{variant}_checkpoint.pt"
|
||||||
|
torch.save(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"variant": variant,
|
||||||
|
"seed": seed,
|
||||||
|
"steps": steps,
|
||||||
|
"absolute_position_encoding": use_pe,
|
||||||
|
"model_state_dict": model.state_dict(),
|
||||||
|
"synthetic_only": True,
|
||||||
|
},
|
||||||
|
checkpoint,
|
||||||
|
)
|
||||||
|
return metric_rows, history_rows, {"checkpoint": str(checkpoint), "final_loss": history[-1]["L_align"]}
|
||||||
|
|
||||||
|
|
||||||
|
def run(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
output_dir = args.output_dir
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
if args.device == "auto":
|
||||||
|
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||||
|
else:
|
||||||
|
device = torch.device(args.device)
|
||||||
|
if device.type == "cuda" and not torch.cuda.is_available():
|
||||||
|
raise RuntimeError("CUDA was requested but is unavailable")
|
||||||
|
variants = [("M3", "M3"), ("M4", "M4_noPE"), ("M4", "M4_sinPE")]
|
||||||
|
all_metrics: list[dict[str, Any]] = []
|
||||||
|
all_history: list[dict[str, Any]] = []
|
||||||
|
model_summaries: list[dict[str, Any]] = []
|
||||||
|
for method, variant in variants:
|
||||||
|
metrics, history, summary = _run_trial(
|
||||||
|
method, variant, device=device, steps=args.steps, seed=args.seed, output_dir=output_dir
|
||||||
|
)
|
||||||
|
all_metrics.extend(metrics)
|
||||||
|
all_history.extend(history)
|
||||||
|
model_summaries.append({"method": method, "variant": variant, **summary})
|
||||||
|
_write_csv(output_dir / "metrics.csv", all_metrics)
|
||||||
|
_write_csv(output_dir / "training_history.csv", all_history)
|
||||||
|
manifest = {
|
||||||
|
"created_utc": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"purpose": "Check whether the existing M3/M4 attention implementation can fit a time band when all inputs encode only synthetic time.",
|
||||||
|
"sample_count": 1,
|
||||||
|
"synthetic_only": True,
|
||||||
|
"seed": args.seed,
|
||||||
|
"device": str(device),
|
||||||
|
"gpu_name": torch.cuda.get_device_name(device) if device.type == "cuda" else None,
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"grid_size": GRID_SIZE,
|
||||||
|
"sigma_normalized_time": SIGMA,
|
||||||
|
"feature_rule": "[t, t^2, 1] per source position; no extracted text/audio/video content is used",
|
||||||
|
"sequence_lengths": MODALITY_LENGTHS,
|
||||||
|
"optimizer": "AdamW",
|
||||||
|
"learning_rate": 1e-3,
|
||||||
|
"steps_per_variant": args.steps,
|
||||||
|
"variants": model_summaries,
|
||||||
|
"interpretation_limits": [
|
||||||
|
"This is an implementation/optimization control only; it does not evaluate real features or alignment accuracy.",
|
||||||
|
"The time-only features deliberately provide source-position information that the real feature inputs may not contain.",
|
||||||
|
],
|
||||||
|
}
|
||||||
|
(output_dir / "run_manifest.json").write_text(
|
||||||
|
json.dumps(manifest, ensure_ascii=False, indent=2, allow_nan=False), encoding="utf-8"
|
||||||
|
)
|
||||||
|
print(f"[all done] wrote synthetic control to {output_dir}", flush=True)
|
||||||
|
return manifest
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
project = Path(__file__).resolve().parents[1]
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument("--steps", type=int, default=1000)
|
||||||
|
parser.add_argument("--seed", type=int, default=SEED)
|
||||||
|
parser.add_argument("--device", choices=("auto", "cpu", "cuda"), default="auto")
|
||||||
|
parser.add_argument(
|
||||||
|
"--output-dir", type=Path, default=project / "outputs/alignment_debug/synthetic"
|
||||||
|
)
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
run(args)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,647 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import math
|
||||||
|
import platform
|
||||||
|
import statistics
|
||||||
|
import time
|
||||||
|
from collections import defaultdict
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Mapping, Sequence
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
from sklearn.model_selection import GroupKFold
|
||||||
|
|
||||||
|
from .compare_methods import (
|
||||||
|
_alignment_rows,
|
||||||
|
_batches,
|
||||||
|
_baseline_output,
|
||||||
|
_collect_representations,
|
||||||
|
_fit_learned_model,
|
||||||
|
_make_figures,
|
||||||
|
_save_alignment,
|
||||||
|
_seed_everything,
|
||||||
|
_write_csv,
|
||||||
|
)
|
||||||
|
from .experiment_data import (
|
||||||
|
FeatureSample,
|
||||||
|
collate_feature_samples,
|
||||||
|
fit_feature_stats,
|
||||||
|
load_feature_samples,
|
||||||
|
)
|
||||||
|
from .experiment_probes import (
|
||||||
|
run_shuffled_alignment_reconstruction_probe,
|
||||||
|
run_within_clip_temporal_retrieval_probe,
|
||||||
|
)
|
||||||
|
from .types import MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
VARIANTS = ("v2_a", "v2_b", "v2_c")
|
||||||
|
METRIC_COLUMNS = {
|
||||||
|
"alignment": (
|
||||||
|
"mvr",
|
||||||
|
"normalized_entropy",
|
||||||
|
"width80_source_positions",
|
||||||
|
"trajectory_span_fraction",
|
||||||
|
"c_row",
|
||||||
|
"c_far",
|
||||||
|
),
|
||||||
|
"retrieval": ("r_at_1", "r_at_3", "mase_slots", "exact_r_at_1"),
|
||||||
|
"reconstruction": (
|
||||||
|
"mae_aligned",
|
||||||
|
"mae_shuffled_mean",
|
||||||
|
"mae_shuffled_std",
|
||||||
|
"gain_align",
|
||||||
|
),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _summarize(
|
||||||
|
rows: Sequence[Mapping[str, Any]], group_keys: Sequence[str], metrics: Sequence[str]
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
grouped: dict[tuple[Any, ...], list[Mapping[str, Any]]] = defaultdict(list)
|
||||||
|
for row in rows:
|
||||||
|
grouped[tuple(row[key] for key in group_keys)].append(row)
|
||||||
|
results = []
|
||||||
|
for key, values in grouped.items():
|
||||||
|
summary: dict[str, Any] = dict(zip(group_keys, key))
|
||||||
|
summary["n"] = len(values)
|
||||||
|
for metric in metrics:
|
||||||
|
numbers = [float(row[metric]) for row in values if row.get(metric) not in (None, "")]
|
||||||
|
numbers = [number for number in numbers if math.isfinite(number)]
|
||||||
|
if numbers:
|
||||||
|
summary[f"{metric}_mean"] = statistics.fmean(numbers)
|
||||||
|
summary[f"{metric}_std"] = statistics.stdev(numbers) if len(numbers) > 1 else 0.0
|
||||||
|
results.append(summary)
|
||||||
|
return results
|
||||||
|
|
||||||
|
|
||||||
|
def _comparison_table(summaries: Mapping[str, Sequence[Mapping[str, Any]]]) -> list[dict[str, Any]]:
|
||||||
|
alignment = {(row["method"], row["modality"]): row for row in summaries["alignment"]}
|
||||||
|
retrieval = {(row["method"], row["direction"]): row for row in summaries["retrieval"]}
|
||||||
|
reconstruction = {
|
||||||
|
(row["method"], row["target_modality"]): row for row in summaries["reconstruction"]
|
||||||
|
}
|
||||||
|
rows = []
|
||||||
|
for method in ("M1", "M2", "M3", "M4"):
|
||||||
|
row: dict[str, Any] = {"method": method}
|
||||||
|
for modality in MODALITIES:
|
||||||
|
metrics = alignment[(method, modality)]
|
||||||
|
for key in ("mvr", "normalized_entropy", "trajectory_span_fraction", "c_row", "c_far"):
|
||||||
|
row[f"{key}_{modality}"] = metrics.get(f"{key}_mean")
|
||||||
|
for direction in ("text_to_audio", "text_to_vision", "audio_to_vision"):
|
||||||
|
metrics = retrieval[(method, direction)]
|
||||||
|
for key in ("r_at_1", "r_at_3", "mase_slots", "exact_r_at_1"):
|
||||||
|
row[f"{key}_{direction}"] = metrics.get(f"{key}_mean")
|
||||||
|
for modality in MODALITIES:
|
||||||
|
metrics = reconstruction[(method, modality)]
|
||||||
|
for key in ("mae_aligned", "mae_shuffled_mean", "gain_align"):
|
||||||
|
row[f"{key}_{modality}"] = metrics.get(f"{key}_mean")
|
||||||
|
rows.append(row)
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def _folds_from_baseline(
|
||||||
|
samples: Sequence[FeatureSample], baseline_splits: Path, requested_folds: int
|
||||||
|
) -> list[tuple[list[FeatureSample], list[FeatureSample], dict[str, Any]]]:
|
||||||
|
by_id = {sample.sample_id: sample for sample in samples}
|
||||||
|
if baseline_splits.is_file():
|
||||||
|
split_rows = json.loads(baseline_splits.read_text(encoding="utf-8"))
|
||||||
|
else:
|
||||||
|
groups = [sample.group_id for sample in samples]
|
||||||
|
splitter = GroupKFold(n_splits=requested_folds)
|
||||||
|
split_rows = []
|
||||||
|
for fold, (train, validation) in enumerate(
|
||||||
|
splitter.split(np.zeros(len(samples)), groups=groups), start=1
|
||||||
|
):
|
||||||
|
split_rows.append(
|
||||||
|
{
|
||||||
|
"fold": fold,
|
||||||
|
"train_sample_ids": [samples[index].sample_id for index in train],
|
||||||
|
"validation_sample_ids": [samples[index].sample_id for index in validation],
|
||||||
|
"train_video_ids": sorted({samples[index].group_id for index in train}),
|
||||||
|
"validation_video_ids": sorted(
|
||||||
|
{samples[index].group_id for index in validation}
|
||||||
|
),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
if len(split_rows) != requested_folds:
|
||||||
|
raise ValueError(
|
||||||
|
f"baseline split file has {len(split_rows)} folds; expected {requested_folds}"
|
||||||
|
)
|
||||||
|
folds = []
|
||||||
|
seen_validation: list[str] = []
|
||||||
|
for row in split_rows:
|
||||||
|
train_ids = row["train_sample_ids"]
|
||||||
|
validation_ids = row["validation_sample_ids"]
|
||||||
|
if set(train_ids) & set(validation_ids):
|
||||||
|
raise ValueError(f"sample leakage in fold {row['fold']}")
|
||||||
|
train_samples = [by_id[sample_id] for sample_id in train_ids]
|
||||||
|
val_samples = [by_id[sample_id] for sample_id in validation_ids]
|
||||||
|
train_groups = {sample.group_id for sample in train_samples}
|
||||||
|
val_groups = {sample.group_id for sample in val_samples}
|
||||||
|
if train_groups & val_groups:
|
||||||
|
raise ValueError(f"video_id leakage in fold {row['fold']}")
|
||||||
|
seen_validation.extend(validation_ids)
|
||||||
|
folds.append((train_samples, val_samples, row))
|
||||||
|
if len(seen_validation) != len(samples) or set(seen_validation) != set(by_id):
|
||||||
|
raise ValueError("baseline folds do not cover the current complete sample manifest")
|
||||||
|
return folds
|
||||||
|
|
||||||
|
|
||||||
|
def _evaluate_fixed_methods(
|
||||||
|
args: argparse.Namespace,
|
||||||
|
samples: Sequence[FeatureSample],
|
||||||
|
folds: Sequence[tuple[list[FeatureSample], list[FeatureSample], dict[str, Any]]],
|
||||||
|
device: torch.device,
|
||||||
|
output_dir: Path,
|
||||||
|
) -> tuple[dict[str, list[dict[str, Any]]], dict[str, dict[str, Mapping[str, np.ndarray]]]]:
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
rows: dict[str, list[dict[str, Any]]] = {
|
||||||
|
"alignment": [],
|
||||||
|
"retrieval": [],
|
||||||
|
"reconstruction": [],
|
||||||
|
}
|
||||||
|
example_weights: dict[str, dict[str, Mapping[str, np.ndarray]]] = defaultdict(dict)
|
||||||
|
sample_by_id = {sample.sample_id: sample for sample in samples}
|
||||||
|
example_id = args.example_id if args.example_id in sample_by_id else samples[0].sample_id
|
||||||
|
|
||||||
|
for fold_index, (train_samples, val_samples, split_row) in enumerate(folds, start=1):
|
||||||
|
stats = fit_feature_stats(train_samples)
|
||||||
|
print(
|
||||||
|
f"[fixed fold {fold_index}/{len(folds)}] train={len(train_samples)} "
|
||||||
|
f"validation={len(val_samples)}; evaluating unchanged M1/M2",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
for method in ("M1", "M2"):
|
||||||
|
train_aligned, _ = _collect_representations(
|
||||||
|
method,
|
||||||
|
train_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
val_aligned, val_weights = _collect_representations(
|
||||||
|
method,
|
||||||
|
val_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
combined = {**train_aligned, **val_aligned}
|
||||||
|
for batch_samples in _batches(
|
||||||
|
val_samples,
|
||||||
|
args.batch_size,
|
||||||
|
shuffle=False,
|
||||||
|
rng=np.random.default_rng(args.seeds[0] + fold_index),
|
||||||
|
):
|
||||||
|
sequences, durations, intervals = collate_feature_samples(
|
||||||
|
batch_samples, stats, device
|
||||||
|
)
|
||||||
|
output = _baseline_output(method, sequences, durations, intervals, args.grid_size)
|
||||||
|
rows["alignment"].extend(
|
||||||
|
_alignment_rows(
|
||||||
|
method,
|
||||||
|
"fixed",
|
||||||
|
fold_index,
|
||||||
|
batch_samples,
|
||||||
|
output,
|
||||||
|
device,
|
||||||
|
args.mvr_epsilon,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
for sample in val_samples:
|
||||||
|
_save_alignment(output_dir, method, "fixed", sample, val_weights[sample.sample_id])
|
||||||
|
if sample.sample_id == example_id:
|
||||||
|
example_weights[sample.sample_id][method] = val_weights[sample.sample_id]
|
||||||
|
|
||||||
|
probe_seed = args.seeds[0] + fold_index * 100
|
||||||
|
rows["retrieval"].extend(
|
||||||
|
{"method": method, "seed": "fixed", "fold": fold_index, **row}
|
||||||
|
for row in run_within_clip_temporal_retrieval_probe(
|
||||||
|
[sample.sample_id for sample in train_samples],
|
||||||
|
[sample.sample_id for sample in val_samples],
|
||||||
|
combined,
|
||||||
|
device=device,
|
||||||
|
seed=probe_seed,
|
||||||
|
epochs=args.retrieval_probe_epochs,
|
||||||
|
batch_size=args.probe_batch_size,
|
||||||
|
tolerance=args.retrieval_tolerance,
|
||||||
|
top_k=3,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
rows["reconstruction"].extend(
|
||||||
|
{"method": method, "seed": "fixed", "fold": fold_index, **row}
|
||||||
|
for row in run_shuffled_alignment_reconstruction_probe(
|
||||||
|
[sample.sample_id for sample in train_samples],
|
||||||
|
[sample.sample_id for sample in val_samples],
|
||||||
|
combined,
|
||||||
|
device=device,
|
||||||
|
seed=probe_seed + 1,
|
||||||
|
ratio=args.mask_ratio,
|
||||||
|
epochs=args.reconstruction_probe_epochs,
|
||||||
|
batch_size=args.probe_batch_size,
|
||||||
|
shuffle_repeats=args.shuffle_repeats,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
for key, values in rows.items():
|
||||||
|
_write_csv(output_dir / f"{key}_metrics.csv", values)
|
||||||
|
(output_dir / "README.md").write_text(
|
||||||
|
"# Fixed M1/M2 reference probes\n\n"
|
||||||
|
"M1 and M2 are recomputed from the same saved video-grouped folds and unchanged. "
|
||||||
|
"The within-clip retrieval projection is fitted on each training fold; held-out candidates "
|
||||||
|
"come only from the same clip. The reconstruction decoder is trained on aligned training "
|
||||||
|
"representations, then compared with a control that shuffles the two non-target streams.\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
return rows, example_weights
|
||||||
|
|
||||||
|
|
||||||
|
def _write_stage_summary(
|
||||||
|
stage_dir: Path,
|
||||||
|
fixed_rows: Mapping[str, Sequence[dict[str, Any]]],
|
||||||
|
learned_rows: Mapping[str, Sequence[dict[str, Any]]],
|
||||||
|
) -> dict[str, list[dict[str, Any]]]:
|
||||||
|
all_rows = {
|
||||||
|
key: [*fixed_rows[key], *learned_rows[key]]
|
||||||
|
for key in ("alignment", "retrieval", "reconstruction")
|
||||||
|
}
|
||||||
|
group_columns = {
|
||||||
|
"alignment": ("method", "modality"),
|
||||||
|
"retrieval": ("method", "direction"),
|
||||||
|
"reconstruction": ("method", "target_modality"),
|
||||||
|
}
|
||||||
|
summaries = {
|
||||||
|
key: _summarize(rows, group_columns[key], METRIC_COLUMNS[key])
|
||||||
|
for key, rows in all_rows.items()
|
||||||
|
}
|
||||||
|
for key, rows in all_rows.items():
|
||||||
|
_write_csv(stage_dir / f"{key}_metrics_with_fixed.csv", rows)
|
||||||
|
_write_csv(stage_dir / f"{key}_summary.csv", summaries[key])
|
||||||
|
_write_csv(stage_dir / "comparison_summary.csv", _comparison_table(summaries))
|
||||||
|
(stage_dir / "summary.json").write_text(
|
||||||
|
json.dumps(summaries, ensure_ascii=False, indent=2, allow_nan=False), encoding="utf-8"
|
||||||
|
)
|
||||||
|
return summaries
|
||||||
|
|
||||||
|
|
||||||
|
def _run_variant(
|
||||||
|
variant: str,
|
||||||
|
args: argparse.Namespace,
|
||||||
|
samples: Sequence[FeatureSample],
|
||||||
|
folds: Sequence[tuple[list[FeatureSample], list[FeatureSample], dict[str, Any]]],
|
||||||
|
fixed_rows: Mapping[str, Sequence[dict[str, Any]]],
|
||||||
|
fixed_examples: Mapping[str, Mapping[str, Mapping[str, np.ndarray]]],
|
||||||
|
device: torch.device,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
start_time = time.time()
|
||||||
|
stage_dir = args.output_dir / variant
|
||||||
|
stage_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
(stage_dir / "splits.json").write_text(
|
||||||
|
json.dumps([row for _, _, row in folds], ensure_ascii=False, indent=2), encoding="utf-8"
|
||||||
|
)
|
||||||
|
learned_rows: dict[str, list[dict[str, Any]]] = {
|
||||||
|
"alignment": [],
|
||||||
|
"retrieval": [],
|
||||||
|
"reconstruction": [],
|
||||||
|
}
|
||||||
|
training_summary: list[dict[str, Any]] = []
|
||||||
|
training_history: list[dict[str, Any]] = []
|
||||||
|
example_weights: dict[str, dict[str, Mapping[str, np.ndarray]]] = defaultdict(dict)
|
||||||
|
sample_by_id = {sample.sample_id: sample for sample in samples}
|
||||||
|
example_id = args.example_id if args.example_id in sample_by_id else samples[0].sample_id
|
||||||
|
example_sample = sample_by_id[example_id]
|
||||||
|
if example_id in fixed_examples:
|
||||||
|
example_weights[example_id].update(fixed_examples[example_id])
|
||||||
|
|
||||||
|
for fold_index, (train_samples, val_samples, _) in enumerate(folds, start=1):
|
||||||
|
stats = fit_feature_stats(train_samples)
|
||||||
|
fold_dir = stage_dir / f"fold_{fold_index:02d}"
|
||||||
|
print(
|
||||||
|
f"[{variant} fold {fold_index}/{len(folds)}] training M3/M4 with "
|
||||||
|
f"{len(train_samples)} train and {len(val_samples)} validation clips",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
for seed in args.seeds:
|
||||||
|
for method in ("M3", "M4"):
|
||||||
|
checkpoint = fold_dir / f"seed_{seed}" / f"{method}.pt"
|
||||||
|
model, info = _fit_learned_model(
|
||||||
|
method,
|
||||||
|
train_samples,
|
||||||
|
val_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
seed=seed + fold_index * 1009,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
hidden_size=args.hidden_size,
|
||||||
|
heads=args.heads,
|
||||||
|
dropout=args.dropout,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
max_epochs=args.epochs,
|
||||||
|
patience=args.patience,
|
||||||
|
learning_rate=args.learning_rate,
|
||||||
|
checkpoint_path=checkpoint,
|
||||||
|
loss_variant=variant,
|
||||||
|
)
|
||||||
|
training_summary.append(
|
||||||
|
{
|
||||||
|
"loss_variant": variant,
|
||||||
|
"method": method,
|
||||||
|
"seed": seed,
|
||||||
|
"fold": fold_index,
|
||||||
|
"best_epoch": info["best_epoch"],
|
||||||
|
"best_validation_objective": info["best_validation_objective"],
|
||||||
|
**{
|
||||||
|
key: value
|
||||||
|
for key, value in info["best_validation_metrics"].items()
|
||||||
|
if key not in {"epoch", "train_total"}
|
||||||
|
},
|
||||||
|
"checkpoint": info["checkpoint"],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
training_history.extend(
|
||||||
|
{
|
||||||
|
"loss_variant": variant,
|
||||||
|
"method": method,
|
||||||
|
"seed": seed,
|
||||||
|
"fold": fold_index,
|
||||||
|
**epoch,
|
||||||
|
}
|
||||||
|
for epoch in info["history"]
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f"[{variant} fold {fold_index}] {method} seed={seed} "
|
||||||
|
f"best_epoch={info['best_epoch']} "
|
||||||
|
f"val={info['best_validation_objective']:.4f} "
|
||||||
|
f"C_row(audio/vision)="
|
||||||
|
f"{info['best_validation_metrics']['validation_c_row_audio']:.3f}/"
|
||||||
|
f"{info['best_validation_metrics']['validation_c_row_vision']:.3f}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
train_aligned, _ = _collect_representations(
|
||||||
|
method,
|
||||||
|
train_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
model=model,
|
||||||
|
)
|
||||||
|
val_aligned, val_weights = _collect_representations(
|
||||||
|
method,
|
||||||
|
val_samples,
|
||||||
|
stats,
|
||||||
|
device=device,
|
||||||
|
grid_size=args.grid_size,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
model=model,
|
||||||
|
)
|
||||||
|
combined = {**train_aligned, **val_aligned}
|
||||||
|
for batch_samples in _batches(
|
||||||
|
val_samples,
|
||||||
|
args.batch_size,
|
||||||
|
shuffle=False,
|
||||||
|
rng=np.random.default_rng(seed + fold_index),
|
||||||
|
):
|
||||||
|
sequences, durations, _ = collate_feature_samples(
|
||||||
|
batch_samples, stats, device
|
||||||
|
)
|
||||||
|
output = model(sequences)
|
||||||
|
learned_rows["alignment"].extend(
|
||||||
|
_alignment_rows(
|
||||||
|
method,
|
||||||
|
str(seed),
|
||||||
|
fold_index,
|
||||||
|
batch_samples,
|
||||||
|
output,
|
||||||
|
device,
|
||||||
|
args.mvr_epsilon,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
for sample in val_samples:
|
||||||
|
_save_alignment(stage_dir, method, str(seed), sample, val_weights[sample.sample_id])
|
||||||
|
if sample.sample_id == example_id and seed == args.seeds[0]:
|
||||||
|
example_weights[sample.sample_id][method] = val_weights[sample.sample_id]
|
||||||
|
|
||||||
|
probe_seed = seed + fold_index * 100 + (3 if method == "M3" else 7)
|
||||||
|
learned_rows["retrieval"].extend(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"seed": seed,
|
||||||
|
"fold": fold_index,
|
||||||
|
**row,
|
||||||
|
}
|
||||||
|
for row in run_within_clip_temporal_retrieval_probe(
|
||||||
|
[sample.sample_id for sample in train_samples],
|
||||||
|
[sample.sample_id for sample in val_samples],
|
||||||
|
combined,
|
||||||
|
device=device,
|
||||||
|
seed=probe_seed,
|
||||||
|
epochs=args.retrieval_probe_epochs,
|
||||||
|
batch_size=args.probe_batch_size,
|
||||||
|
tolerance=args.retrieval_tolerance,
|
||||||
|
top_k=3,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
learned_rows["reconstruction"].extend(
|
||||||
|
{
|
||||||
|
"method": method,
|
||||||
|
"seed": seed,
|
||||||
|
"fold": fold_index,
|
||||||
|
**row,
|
||||||
|
}
|
||||||
|
for row in run_shuffled_alignment_reconstruction_probe(
|
||||||
|
[sample.sample_id for sample in train_samples],
|
||||||
|
[sample.sample_id for sample in val_samples],
|
||||||
|
combined,
|
||||||
|
device=device,
|
||||||
|
seed=probe_seed + 1,
|
||||||
|
ratio=args.mask_ratio,
|
||||||
|
epochs=args.reconstruction_probe_epochs,
|
||||||
|
batch_size=args.probe_batch_size,
|
||||||
|
shuffle_repeats=args.shuffle_repeats,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
del model
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
|
||||||
|
_write_csv(stage_dir / "training_summary.csv", training_summary)
|
||||||
|
_write_csv(stage_dir / "training_history.csv", training_history)
|
||||||
|
for key, rows in learned_rows.items():
|
||||||
|
_write_csv(stage_dir / f"{key}_metrics_learned_only.csv", rows)
|
||||||
|
|
||||||
|
summaries = _write_stage_summary(stage_dir, fixed_rows, learned_rows)
|
||||||
|
_write_csv(stage_dir / "training_summary.csv", training_summary)
|
||||||
|
_write_csv(stage_dir / "training_history.csv", training_history)
|
||||||
|
if example_id in example_weights and set(example_weights[example_id]) == {"M1", "M2", "M3", "M4"}:
|
||||||
|
_make_figures(stage_dir, example_id, example_sample, example_weights[example_id], args.grid_size)
|
||||||
|
manifest = {
|
||||||
|
"variant": variant,
|
||||||
|
"loss_coefficients": {
|
||||||
|
"lambda_reconstruction": 1.0,
|
||||||
|
"lambda_contrastive": 1.0,
|
||||||
|
"lambda_monotonicity": 0.1,
|
||||||
|
"lambda_span": 5.0,
|
||||||
|
"lambda_diversity": 0.5 if variant in {"v2_b", "v2_c"} else 0.0,
|
||||||
|
"lambda_band": 10.0 if variant == "v2_c" else 0.0,
|
||||||
|
"coverage_floor": 0.7,
|
||||||
|
"diversity_slot_separation": 6,
|
||||||
|
"band_margin": 0.1,
|
||||||
|
},
|
||||||
|
"sample_count": len(samples),
|
||||||
|
"video_group_folds": len(folds),
|
||||||
|
"seeds": args.seeds,
|
||||||
|
"device": str(device),
|
||||||
|
"gpu_name": torch.cuda.get_device_name(device) if device.type == "cuda" else None,
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"elapsed_seconds": time.time() - start_time,
|
||||||
|
"probes": {
|
||||||
|
"within_clip_retrieval_tolerance_slots": args.retrieval_tolerance,
|
||||||
|
"retrieval_top_k": 3,
|
||||||
|
"shuffled_reconstruction_repeats": args.shuffle_repeats,
|
||||||
|
"masked_block_ratio": args.mask_ratio,
|
||||||
|
},
|
||||||
|
"limits": [
|
||||||
|
"Retrieval projections are fitted on training-fold grid-slot positives; test candidates are restricted to the same held-out clip.",
|
||||||
|
"The reconstruction control shuffles the two non-target modality slot streams and preserves the target stream.",
|
||||||
|
"No human event timestamps are available, so temporal probes do not replace manual annotation.",
|
||||||
|
"Slot-regularization losses impose weak temporal structure and must be interpreted alongside the unregularized M1/M2 reference.",
|
||||||
|
],
|
||||||
|
}
|
||||||
|
(stage_dir / "run_manifest.json").write_text(
|
||||||
|
json.dumps(manifest, ensure_ascii=False, indent=2, allow_nan=False), encoding="utf-8"
|
||||||
|
)
|
||||||
|
(stage_dir / "README.md").write_text(
|
||||||
|
f"# {variant} M3/M4 alignment variant\n\n"
|
||||||
|
"M1/M2 rows in the comparison files are fixed references recomputed on the original grouped folds. "
|
||||||
|
"Only M3/M4 training losses changed. See `training_history.csv` for per-epoch loss components and row-collapse scores. "
|
||||||
|
"`retrieval_metrics_learned_only.csv` restricts candidates to the same clip. "
|
||||||
|
"`reconstruction_metrics_learned_only.csv` contrasts aligned and shuffled non-target streams. "
|
||||||
|
"The experiment does not include RoPE, Gaussian bias, or latent-length changes.\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f"[{variant} done] elapsed={manifest['elapsed_seconds']:.1f}s "
|
||||||
|
f"output={stage_dir}; summary rows={sum(len(rows) for rows in summaries.values())}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
return manifest
|
||||||
|
|
||||||
|
|
||||||
|
def run(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
start_time = time.time()
|
||||||
|
_seed_everything(args.seeds[0])
|
||||||
|
if args.device == "auto":
|
||||||
|
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
|
||||||
|
else:
|
||||||
|
device = torch.device(args.device)
|
||||||
|
if device.type == "cuda" and not torch.cuda.is_available():
|
||||||
|
raise RuntimeError("CUDA was requested but is not available")
|
||||||
|
samples = load_feature_samples(args.feature_dir, args.manifest)
|
||||||
|
if len(samples) != args.expected_samples:
|
||||||
|
raise ValueError(f"expected {args.expected_samples} samples, found {len(samples)}")
|
||||||
|
folds = _folds_from_baseline(
|
||||||
|
samples, args.baseline_dir / "splits.json", args.folds
|
||||||
|
)
|
||||||
|
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
print(
|
||||||
|
f"[start] samples={len(samples)} folds={len(folds)} seeds={args.seeds} "
|
||||||
|
f"device={device} variants={','.join(VARIANTS)}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
fixed_rows, fixed_examples = _evaluate_fixed_methods(
|
||||||
|
args,
|
||||||
|
samples,
|
||||||
|
folds,
|
||||||
|
device,
|
||||||
|
args.output_dir / "fixed_baselines",
|
||||||
|
)
|
||||||
|
stage_manifests = []
|
||||||
|
for variant in VARIANTS:
|
||||||
|
stage_manifests.append(
|
||||||
|
_run_variant(
|
||||||
|
variant,
|
||||||
|
args,
|
||||||
|
samples,
|
||||||
|
folds,
|
||||||
|
fixed_rows,
|
||||||
|
fixed_examples,
|
||||||
|
device,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
manifest = {
|
||||||
|
"created_utc": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"sample_count": len(samples),
|
||||||
|
"group_count": len({sample.group_id for sample in samples}),
|
||||||
|
"folds": len(folds),
|
||||||
|
"seeds": args.seeds,
|
||||||
|
"device": str(device),
|
||||||
|
"gpu_name": torch.cuda.get_device_name(device) if device.type == "cuda" else None,
|
||||||
|
"python": platform.python_version(),
|
||||||
|
"torch": torch.__version__,
|
||||||
|
"feature_dir": str(args.feature_dir),
|
||||||
|
"baseline_dir": str(args.baseline_dir),
|
||||||
|
"variants": stage_manifests,
|
||||||
|
"elapsed_seconds": time.time() - start_time,
|
||||||
|
"fixed_methods_unchanged": ["M1", "M2"],
|
||||||
|
"feature_extraction_changed": False,
|
||||||
|
}
|
||||||
|
(args.output_dir / "run_manifest.json").write_text(
|
||||||
|
json.dumps(manifest, ensure_ascii=False, indent=2, allow_nan=False), encoding="utf-8"
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
f"[all done] elapsed={manifest['elapsed_seconds']:.1f}s output={args.output_dir}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
return manifest
|
||||||
|
|
||||||
|
|
||||||
|
def build_parser() -> argparse.ArgumentParser:
|
||||||
|
project_dir = Path(__file__).resolve().parents[1]
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Train staged M3/M4 alignment-loss variants and stronger temporal probes."
|
||||||
|
)
|
||||||
|
parser.add_argument("--feature-dir", type=Path, default=project_dir / "outputs/q1_features/features")
|
||||||
|
parser.add_argument("--manifest", type=Path, default=project_dir / "outputs/audit/manifest.csv")
|
||||||
|
parser.add_argument("--baseline-dir", type=Path, default=project_dir / "outputs/method_comparison")
|
||||||
|
parser.add_argument("--output-dir", type=Path, default=project_dir / "outputs/alignment_v2")
|
||||||
|
parser.add_argument("--device", default="auto", help="auto, cpu, or a torch device such as cuda:0")
|
||||||
|
parser.add_argument("--expected-samples", type=int, default=100)
|
||||||
|
parser.add_argument("--folds", type=int, default=5)
|
||||||
|
parser.add_argument("--seeds", type=int, nargs="+", default=[42, 3407, 2026])
|
||||||
|
parser.add_argument("--grid-size", type=int, default=50)
|
||||||
|
parser.add_argument("--hidden-size", type=int, default=128)
|
||||||
|
parser.add_argument("--heads", type=int, default=4)
|
||||||
|
parser.add_argument("--dropout", type=float, default=0.1)
|
||||||
|
parser.add_argument("--batch-size", type=int, default=8)
|
||||||
|
parser.add_argument("--probe-batch-size", type=int, default=16)
|
||||||
|
parser.add_argument("--epochs", type=int, default=50)
|
||||||
|
parser.add_argument("--patience", type=int, default=8)
|
||||||
|
parser.add_argument("--learning-rate", type=float, default=1e-4)
|
||||||
|
parser.add_argument("--mask-ratio", type=float, default=0.2)
|
||||||
|
parser.add_argument("--retrieval-probe-epochs", type=int, default=20)
|
||||||
|
parser.add_argument("--reconstruction-probe-epochs", type=int, default=25)
|
||||||
|
parser.add_argument("--retrieval-tolerance", type=int, default=1)
|
||||||
|
parser.add_argument("--shuffle-repeats", type=int, default=5)
|
||||||
|
parser.add_argument("--mvr-epsilon", type=float, default=0.02)
|
||||||
|
parser.add_argument("--example-id", default="-tPCytz4rww/12")
|
||||||
|
return parser
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
manifest = run(args)
|
||||||
|
print(json.dumps(manifest, ensure_ascii=False, indent=2))
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,353 @@
|
|||||||
|
"""Evaluate TSFA attention maps using identical raw source content for every method.
|
||||||
|
|
||||||
|
This probe pools the same train-fold-standardized BERT, audio, and DeiT features
|
||||||
|
with each method's alignment matrix. It excludes native model value/output
|
||||||
|
projections from the representation being scored.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import csv
|
||||||
|
import json
|
||||||
|
import shutil
|
||||||
|
import time
|
||||||
|
from datetime import datetime, timezone
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Mapping, Sequence
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import torch
|
||||||
|
|
||||||
|
from .correspondence_eval import _cluster_bootstrap, _write_csv
|
||||||
|
from .experiment_data import (
|
||||||
|
FeatureSample,
|
||||||
|
collate_feature_samples,
|
||||||
|
fit_feature_stats,
|
||||||
|
load_feature_samples,
|
||||||
|
)
|
||||||
|
from .m4_shared_latent_eval import _bootstrap_summary
|
||||||
|
from .tsfa_experiment import (
|
||||||
|
ALL_METHODS,
|
||||||
|
BASELINE_VARIANTS,
|
||||||
|
GRID_SIZE,
|
||||||
|
TSFA_VARIANTS,
|
||||||
|
_ablation_summary,
|
||||||
|
_collect_fold_features,
|
||||||
|
_content_summary,
|
||||||
|
_evaluate_fixed_projector,
|
||||||
|
_fit_method_probe,
|
||||||
|
_generate_tsfa_outputs,
|
||||||
|
_load_semantic_checkpoint,
|
||||||
|
build_parser,
|
||||||
|
)
|
||||||
|
from .types import MODALITIES
|
||||||
|
|
||||||
|
|
||||||
|
def _pool_same_raw_features(
|
||||||
|
samples: Sequence[FeatureSample],
|
||||||
|
feature_stats: Any,
|
||||||
|
weights_by_id: Mapping[str, Mapping[str, np.ndarray]],
|
||||||
|
device: torch.device,
|
||||||
|
batch_size: int,
|
||||||
|
) -> dict[str, dict[str, np.ndarray]]:
|
||||||
|
content: dict[str, dict[str, np.ndarray]] = {}
|
||||||
|
with torch.no_grad():
|
||||||
|
for start in range(0, len(samples), batch_size):
|
||||||
|
batch_samples = list(samples[start : start + batch_size])
|
||||||
|
sequences, _, _ = collate_feature_samples(batch_samples, feature_stats, device)
|
||||||
|
for index, sample in enumerate(batch_samples):
|
||||||
|
content[sample.sample_id] = {}
|
||||||
|
for modality in MODALITIES:
|
||||||
|
length = len(sample.features[modality])
|
||||||
|
weights = torch.as_tensor(
|
||||||
|
weights_by_id[sample.sample_id][modality],
|
||||||
|
dtype=torch.float32,
|
||||||
|
device=device,
|
||||||
|
)
|
||||||
|
if weights.shape != (GRID_SIZE, length):
|
||||||
|
raise ValueError(f"unexpected alignment shape for {sample.sample_id}/{modality}")
|
||||||
|
source = sequences[modality].features[index, :length]
|
||||||
|
content[sample.sample_id][modality] = (
|
||||||
|
(weights @ source).cpu().numpy().astype(np.float32, copy=False)
|
||||||
|
)
|
||||||
|
return content
|
||||||
|
|
||||||
|
|
||||||
|
def _paired_summary(rows: Sequence[Mapping[str, Any]], seed: int) -> list[dict[str, Any]]:
|
||||||
|
by_key = {(row["method"], row["sample_id"]): row for row in rows}
|
||||||
|
sample_ids = sorted({row["sample_id"] for row in rows})
|
||||||
|
contrasts = (
|
||||||
|
("TSFA-main", "M4_sourceTime"),
|
||||||
|
("TSFA-main", "TSFA-random"),
|
||||||
|
("TSFA-main", "TSFA-global"),
|
||||||
|
("TSFA-multiply", "TSFA-main"),
|
||||||
|
("TSFA-main", "M3_noSourceTime"),
|
||||||
|
)
|
||||||
|
output = []
|
||||||
|
for contrast_index, (left, right) in enumerate(contrasts):
|
||||||
|
for metric_index, metric in enumerate((
|
||||||
|
"content_auc_mean", "canonical_pairwise_time_mae_mean"
|
||||||
|
)):
|
||||||
|
differences = []
|
||||||
|
for sample_id in sample_ids:
|
||||||
|
left_row = by_key[(left, sample_id)]
|
||||||
|
right_row = by_key[(right, sample_id)]
|
||||||
|
if left_row["video_id"] != right_row["video_id"]:
|
||||||
|
raise ValueError(f"video group mismatch for {sample_id}")
|
||||||
|
differences.append({
|
||||||
|
"video_id": left_row["video_id"],
|
||||||
|
"difference": float(left_row[metric]) - float(right_row[metric]),
|
||||||
|
})
|
||||||
|
mean, low, high, groups = _cluster_bootstrap(
|
||||||
|
differences,
|
||||||
|
"difference",
|
||||||
|
seed=seed + contrast_index * 101 + metric_index,
|
||||||
|
repetitions=2000,
|
||||||
|
)
|
||||||
|
output.append({
|
||||||
|
"left_method": left,
|
||||||
|
"right_method": right,
|
||||||
|
"metric": metric,
|
||||||
|
"clip_count": len(differences),
|
||||||
|
"video_id_count": groups,
|
||||||
|
"left_minus_right_video_macro_mean": mean,
|
||||||
|
"ci95_low": low,
|
||||||
|
"ci95_high": high,
|
||||||
|
})
|
||||||
|
return output
|
||||||
|
|
||||||
|
|
||||||
|
def _temporal_diagnostic_rows(
|
||||||
|
method: str,
|
||||||
|
fold: int,
|
||||||
|
sample: FeatureSample,
|
||||||
|
weights: Mapping[str, np.ndarray],
|
||||||
|
draw: int = 0,
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
output = []
|
||||||
|
for modality in MODALITIES:
|
||||||
|
valid = np.asarray(sample.valid[modality], dtype=bool)
|
||||||
|
attention = np.asarray(weights[modality], dtype=np.float64)[:, valid]
|
||||||
|
times = np.asarray(sample.times[modality], dtype=np.float64)[valid] / max(sample.duration_s, 1e-8)
|
||||||
|
centers = attention @ times
|
||||||
|
backward = centers[:-1] - centers[1:]
|
||||||
|
entropy = -(attention * np.log(np.maximum(attention, 1e-12))).sum(axis=1)
|
||||||
|
output.append({
|
||||||
|
"method": method,
|
||||||
|
"fold": fold,
|
||||||
|
"sample_id": sample.sample_id,
|
||||||
|
"video_id": sample.group_id,
|
||||||
|
"draw": draw,
|
||||||
|
"modality": modality,
|
||||||
|
"mvr_epsilon_0_01": float(np.mean(backward > 0.01)),
|
||||||
|
"mvr_epsilon_0_02": float(np.mean(backward > 0.02)),
|
||||||
|
"mvr_epsilon_0_05": float(np.mean(backward > 0.05)),
|
||||||
|
"time_span_ratio": float(centers[-1] - centers[0]),
|
||||||
|
"normalized_entropy_mean": float(entropy.mean() / max(np.log(attention.shape[1]), 1e-12)),
|
||||||
|
"source_coverage_rate": float(np.mean(attention.sum(axis=0) > 1e-12)),
|
||||||
|
})
|
||||||
|
return output
|
||||||
|
|
||||||
|
|
||||||
|
def run(args: Any) -> None:
|
||||||
|
started = time.time()
|
||||||
|
device = torch.device(
|
||||||
|
("cuda" if torch.cuda.is_available() else "cpu") if args.device == "auto" else args.device
|
||||||
|
)
|
||||||
|
if device.type == "cuda" and not torch.cuda.is_available():
|
||||||
|
raise RuntimeError("CUDA is unavailable")
|
||||||
|
samples = load_feature_samples(args.feature_dir, args.manifest)
|
||||||
|
samples_by_id = {sample.sample_id: sample for sample in samples}
|
||||||
|
splits = json.loads(args.splits.read_text(encoding="utf-8"))
|
||||||
|
if len(splits) != 5:
|
||||||
|
raise ValueError("expected five grouped folds")
|
||||||
|
store = torch.load(args.output_dir / "probe_checkpoints.pt", map_location="cpu", weights_only=False)
|
||||||
|
content_rows: list[dict[str, Any]] = []
|
||||||
|
curve_rows: list[dict[str, Any]] = []
|
||||||
|
temporal_diagnostic_rows: list[dict[str, Any]] = []
|
||||||
|
history_rows: list[dict[str, Any]] = []
|
||||||
|
probe_store: dict[str, Any] = {}
|
||||||
|
heldout_ids = []
|
||||||
|
|
||||||
|
for split in splits:
|
||||||
|
fold = int(split["fold"])
|
||||||
|
train = [samples_by_id[sample_id] for sample_id in split["train_sample_ids"]]
|
||||||
|
validation = [samples_by_id[sample_id] for sample_id in split["validation_sample_ids"]]
|
||||||
|
if {sample.group_id for sample in train} & {sample.group_id for sample in validation}:
|
||||||
|
raise ValueError(f"video_id leakage in fold {fold}")
|
||||||
|
heldout_ids.extend(sample.sample_id for sample in validation)
|
||||||
|
feature_stats = fit_feature_stats(train)
|
||||||
|
_, baseline_weights, temporal_by_id = _collect_fold_features(
|
||||||
|
fold=fold,
|
||||||
|
train_samples=train,
|
||||||
|
validation_samples=validation,
|
||||||
|
feature_stats=feature_stats,
|
||||||
|
checkpoint_root=args.checkpoint_root,
|
||||||
|
device=device,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
branch = _load_semantic_checkpoint(store, fold, device)
|
||||||
|
all_samples = [*train, *validation]
|
||||||
|
all_ids = [sample.sample_id for sample in all_samples]
|
||||||
|
fold_weights = {method: baseline_weights[method] for method in BASELINE_VARIANTS}
|
||||||
|
for method in TSFA_VARIANTS:
|
||||||
|
_, weights, _ = _generate_tsfa_outputs(
|
||||||
|
method=method,
|
||||||
|
fold=fold,
|
||||||
|
sample_ids=all_ids,
|
||||||
|
samples_by_id=samples_by_id,
|
||||||
|
temporal_by_id=temporal_by_id,
|
||||||
|
branch=branch,
|
||||||
|
device=device,
|
||||||
|
delta=args.delta,
|
||||||
|
seed=args.seed,
|
||||||
|
draw=0,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
fold_weights[method] = weights
|
||||||
|
|
||||||
|
for method in ALL_METHODS:
|
||||||
|
pooled = _pool_same_raw_features(
|
||||||
|
all_samples, feature_stats, fold_weights[method], device, args.batch_size
|
||||||
|
)
|
||||||
|
metrics, curves, _ = _fit_method_probe(
|
||||||
|
method=method,
|
||||||
|
fold=fold,
|
||||||
|
train_samples=train,
|
||||||
|
validation_samples=validation,
|
||||||
|
content_by_id=pooled,
|
||||||
|
device=device,
|
||||||
|
args=args,
|
||||||
|
history_rows=history_rows,
|
||||||
|
checkpoint_store=probe_store,
|
||||||
|
)
|
||||||
|
content_rows.extend(metrics)
|
||||||
|
curve_rows.extend(curves)
|
||||||
|
for sample in validation:
|
||||||
|
temporal_diagnostic_rows.extend(_temporal_diagnostic_rows(
|
||||||
|
method, fold, sample, fold_weights[method][sample.sample_id]
|
||||||
|
))
|
||||||
|
if method == "TSFA-random":
|
||||||
|
for draw in range(1, args.random_window_repeats):
|
||||||
|
_, random_weights, _ = _generate_tsfa_outputs(
|
||||||
|
method=method,
|
||||||
|
fold=fold,
|
||||||
|
sample_ids=[sample.sample_id for sample in validation],
|
||||||
|
samples_by_id=samples_by_id,
|
||||||
|
temporal_by_id=temporal_by_id,
|
||||||
|
branch=branch,
|
||||||
|
device=device,
|
||||||
|
delta=args.delta,
|
||||||
|
seed=args.seed,
|
||||||
|
draw=draw,
|
||||||
|
batch_size=args.batch_size,
|
||||||
|
)
|
||||||
|
random_pooled = _pool_same_raw_features(
|
||||||
|
validation, feature_stats, random_weights, device, args.batch_size
|
||||||
|
)
|
||||||
|
repeated_metrics, repeated_curves, _ = _evaluate_fixed_projector(
|
||||||
|
method=method,
|
||||||
|
fold=fold,
|
||||||
|
validation_samples=validation,
|
||||||
|
content_by_id=random_pooled,
|
||||||
|
projector_state=probe_store[
|
||||||
|
f"fold_{fold:02d}/{method}/correspondence_probe"
|
||||||
|
]["state_dict"],
|
||||||
|
device=device,
|
||||||
|
)
|
||||||
|
content_rows.extend({**row, "draw": draw} for row in repeated_metrics)
|
||||||
|
curve_rows.extend({**row, "draw": draw} for row in repeated_curves)
|
||||||
|
for sample in validation:
|
||||||
|
temporal_diagnostic_rows.extend(_temporal_diagnostic_rows(
|
||||||
|
method, fold, sample, random_weights[sample.sample_id], draw
|
||||||
|
))
|
||||||
|
print(f"[TSFA alignment-only fold {fold}] heldout={len(validation)}", flush=True)
|
||||||
|
del branch, baseline_weights, temporal_by_id, fold_weights
|
||||||
|
if device.type == "cuda":
|
||||||
|
torch.cuda.empty_cache()
|
||||||
|
|
||||||
|
if len(heldout_ids) != 100 or len(set(heldout_ids)) != 100:
|
||||||
|
raise ValueError("held-out fold coverage is not exactly 100 distinct samples")
|
||||||
|
output_dir = args.output_dir
|
||||||
|
content_summary, curve_summary = _content_summary(content_rows, curve_rows, args.seed + 901)
|
||||||
|
with (output_dir / "temporal_metrics_by_clip.csv").open(
|
||||||
|
newline="", encoding="utf-8-sig"
|
||||||
|
) as handle:
|
||||||
|
temporal_rows = list(csv.DictReader(handle))
|
||||||
|
ablation_summary, ablation_by_clip = _ablation_summary(
|
||||||
|
content_rows, temporal_rows, args.seed + 902
|
||||||
|
)
|
||||||
|
paired = _paired_summary(ablation_by_clip, args.seed + 903)
|
||||||
|
temporal_diagnostic_summary = _bootstrap_summary(
|
||||||
|
temporal_diagnostic_rows,
|
||||||
|
("method", "modality"),
|
||||||
|
(
|
||||||
|
"mvr_epsilon_0_01", "mvr_epsilon_0_02", "mvr_epsilon_0_05",
|
||||||
|
"time_span_ratio", "normalized_entropy_mean", "source_coverage_rate",
|
||||||
|
),
|
||||||
|
seed=args.seed + 904,
|
||||||
|
)
|
||||||
|
_write_csv(output_dir / "alignment_only_content_by_clip.csv", content_rows)
|
||||||
|
_write_csv(output_dir / "alignment_only_content_summary.csv", content_summary)
|
||||||
|
_write_csv(output_dir / "alignment_only_shift_curve_summary.csv", curve_summary)
|
||||||
|
_write_csv(output_dir / "alignment_only_ablation_by_clip.csv", ablation_by_clip)
|
||||||
|
_write_csv(output_dir / "alignment_only_ablation_summary.csv", ablation_summary)
|
||||||
|
_write_csv(output_dir / "alignment_only_paired_contrasts.csv", paired)
|
||||||
|
_write_csv(output_dir / "alignment_only_probe_training_history.csv", history_rows)
|
||||||
|
_write_csv(output_dir / "tsfa_temporal_diagnostics_by_clip.csv", temporal_diagnostic_rows)
|
||||||
|
_write_csv(output_dir / "tsfa_temporal_diagnostics_summary.csv", temporal_diagnostic_summary)
|
||||||
|
torch.save(probe_store, output_dir / "alignment_only_probe_checkpoints.pt")
|
||||||
|
run_manifest = {
|
||||||
|
"created_utc": datetime.now(timezone.utc).isoformat(),
|
||||||
|
"experiment": "TSFA shared raw-content alignment-only probe",
|
||||||
|
"sample_count": len(samples),
|
||||||
|
"heldout_count": len(heldout_ids),
|
||||||
|
"video_id_count": len({sample.group_id for sample in samples}),
|
||||||
|
"fold_count": len(splits),
|
||||||
|
"seed": args.seed,
|
||||||
|
"delta": args.delta,
|
||||||
|
"probe_epochs": args.probe_epochs,
|
||||||
|
"random_window_repeats": args.random_window_repeats,
|
||||||
|
"feature_protocol": "Within each fold, feature normalization is fitted on 80 training clips. For every method and modality, the same normalized raw source feature matrix is pooled using that method's A^m; native M3/M4/TSFA value and output projections are excluded from scored features.",
|
||||||
|
"probe_protocol": "One 64-dimensional linear projector per modality, trained on 80 training clips with the same within-clip InfoNCE protocol and identical fold seed across methods. Scores use only 20 held-out clips per fold.",
|
||||||
|
"interpretation_limit": "A same-slot positive is a timestamp/slot convention, not independently annotated semantic ground truth. Source-time attention can still encode time through selected raw values.",
|
||||||
|
"device": str(device),
|
||||||
|
"elapsed_seconds": time.time() - started,
|
||||||
|
}
|
||||||
|
(output_dir / "alignment_only_run_manifest.json").write_text(
|
||||||
|
json.dumps(run_manifest, ensure_ascii=False, indent=2), encoding="utf-8"
|
||||||
|
)
|
||||||
|
bundle = output_dir / "report_bundle"
|
||||||
|
for name in (
|
||||||
|
"alignment_only_content_summary.csv", "alignment_only_ablation_summary.csv",
|
||||||
|
"alignment_only_paired_contrasts.csv", "alignment_only_run_manifest.json",
|
||||||
|
"tsfa_temporal_diagnostics_summary.csv",
|
||||||
|
):
|
||||||
|
shutil.copy2(output_dir / name, bundle / name)
|
||||||
|
bundle_readme = bundle / "README.md"
|
||||||
|
note = (
|
||||||
|
"\nThe `alignment_only_*` summaries pool identical train-fold-standardized "
|
||||||
|
"raw source features through each method's alignment matrix. They isolate "
|
||||||
|
"source selection from native model value/output projections.\n"
|
||||||
|
)
|
||||||
|
current = bundle_readme.read_text(encoding="utf-8")
|
||||||
|
if "The `alignment_only_*` summaries" not in current:
|
||||||
|
bundle_readme.write_text(current + note, encoding="utf-8")
|
||||||
|
print(
|
||||||
|
f"[TSFA alignment-only complete] samples={len(heldout_ids)} "
|
||||||
|
f"elapsed={run_manifest['elapsed_seconds']:.1f}s output={output_dir}",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
args = build_parser().parse_args()
|
||||||
|
if args.finalize_existing:
|
||||||
|
raise ValueError("--finalize-existing belongs to q1.tsfa_experiment")
|
||||||
|
if not 0 < args.delta <= 1:
|
||||||
|
raise ValueError("--delta must be in (0,1]")
|
||||||
|
run(args)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,86 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from typing import Mapping
|
||||||
|
|
||||||
|
import torch
|
||||||
|
from torch import Tensor
|
||||||
|
|
||||||
|
MODALITIES = ("text", "audio", "vision")
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class SequenceBatch:
|
||||||
|
"""A padded batch of timed feature sequences.
|
||||||
|
|
||||||
|
``features`` has shape ``[batch, length, dimension]``; ``times`` and
|
||||||
|
``valid`` have shape ``[batch, length]``. Times are seconds from the start
|
||||||
|
of each clip. Padding positions must be false in ``valid``.
|
||||||
|
"""
|
||||||
|
|
||||||
|
features: Tensor
|
||||||
|
times: Tensor
|
||||||
|
valid: Tensor
|
||||||
|
|
||||||
|
def __post_init__(self) -> None:
|
||||||
|
if self.features.ndim != 3:
|
||||||
|
raise ValueError("features must have shape [batch, length, dimension]")
|
||||||
|
expected = self.features.shape[:2]
|
||||||
|
if tuple(self.times.shape) != tuple(expected):
|
||||||
|
raise ValueError("times must match the batch and sequence dimensions")
|
||||||
|
if tuple(self.valid.shape) != tuple(expected):
|
||||||
|
raise ValueError("valid must match the batch and sequence dimensions")
|
||||||
|
if self.valid.dtype != torch.bool:
|
||||||
|
raise TypeError("valid must be a boolean tensor")
|
||||||
|
if self.features.device != self.times.device or self.features.device != self.valid.device:
|
||||||
|
raise ValueError("features, times, and valid must be on the same device")
|
||||||
|
if not bool(self.valid.any(dim=1).all()):
|
||||||
|
raise ValueError("every sample must contain at least one valid feature")
|
||||||
|
if not bool(torch.isfinite(self.features[self.valid]).all()):
|
||||||
|
raise ValueError("valid features must be finite")
|
||||||
|
if not bool(torch.isfinite(self.times[self.valid]).all()):
|
||||||
|
raise ValueError("valid timestamps must be finite")
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class AlignmentOutput:
|
||||||
|
"""Unified alignment result for Text, Audio, and Vision.
|
||||||
|
|
||||||
|
Each ``weights[m]`` is a row-stochastic matrix with shape ``[B, K, L_m]``.
|
||||||
|
``aligned[m]`` is the resulting common-grid representation ``[B, K, D_m]``.
|
||||||
|
"""
|
||||||
|
|
||||||
|
weights: Mapping[str, Tensor]
|
||||||
|
aligned: Mapping[str, Tensor]
|
||||||
|
fallback_rows: Mapping[str, int] | None = None
|
||||||
|
|
||||||
|
def validate(self, valid: Mapping[str, Tensor] | None = None, atol: float = 1e-4) -> None:
|
||||||
|
if set(self.weights) != set(MODALITIES) or set(self.aligned) != set(MODALITIES):
|
||||||
|
raise ValueError(f"weights and aligned must contain exactly {MODALITIES}")
|
||||||
|
|
||||||
|
batch_size: int | None = None
|
||||||
|
grid_size: int | None = None
|
||||||
|
for modality in MODALITIES:
|
||||||
|
matrix = self.weights[modality]
|
||||||
|
values = self.aligned[modality]
|
||||||
|
if matrix.ndim != 3 or values.ndim != 3:
|
||||||
|
raise ValueError(f"{modality} alignment and values must be rank 3")
|
||||||
|
if matrix.shape[:2] != values.shape[:2]:
|
||||||
|
raise ValueError(f"{modality} weights and aligned values disagree on [B, K]")
|
||||||
|
if batch_size is None:
|
||||||
|
batch_size, grid_size = matrix.shape[:2]
|
||||||
|
elif matrix.shape[0] != batch_size or matrix.shape[1] != grid_size:
|
||||||
|
raise ValueError("all modalities must use the same batch and grid sizes")
|
||||||
|
if not bool(torch.isfinite(matrix).all()) or bool((matrix < -atol).any()):
|
||||||
|
raise ValueError(f"{modality} weights must be finite and non-negative")
|
||||||
|
if not torch.allclose(
|
||||||
|
matrix.sum(dim=-1),
|
||||||
|
torch.ones_like(matrix.sum(dim=-1)),
|
||||||
|
atol=atol,
|
||||||
|
rtol=0,
|
||||||
|
):
|
||||||
|
raise ValueError(f"{modality} alignment rows must sum to one")
|
||||||
|
if valid is not None:
|
||||||
|
allowed = valid[modality][:, None, :]
|
||||||
|
if bool((matrix.masked_select(~allowed.expand_as(matrix)).abs() > atol).any()):
|
||||||
|
raise ValueError(f"{modality} alignment assigns weight to padding")
|
||||||
Generated
+1765
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,136 @@
|
|||||||
|
# Q1 特征提取方法说明(零基础版)
|
||||||
|
|
||||||
|
这份说明介绍本次已经完成的 100 条视频处理过程。目标是把文字、声音和画面变成电脑能处理的数字,并让这三种数字都能回到原视频的时间位置。
|
||||||
|
|
||||||
|
## 先用一句话理解
|
||||||
|
|
||||||
|
把每条视频想成一段带字幕的短片:
|
||||||
|
|
||||||
|
- 把字幕切成一个个词,为每个词生成一组数字;
|
||||||
|
- 把声音切成许多很短的小段,为每段记录音高、响度等声音特征;
|
||||||
|
- 每隔一小段时间取一帧画面,为每帧生成一组画面数字;
|
||||||
|
- 再估计每个字幕词在声音里出现的时间,把同一时间附近的声音和画面对到一起。
|
||||||
|
|
||||||
|
这些数字就是“特征”。它们不是视频本身,也不是情感结论,而是模型后续可以读取的数字表示。
|
||||||
|
|
||||||
|
## 整体流程
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
A[100条原始视频 + 原始字幕] --> B[文字:每个词生成向量]
|
||||||
|
A --> C[声音:提取逐帧声学特征]
|
||||||
|
A --> D[画面:每秒抽取约5帧并生成特征]
|
||||||
|
B --> E[用字幕和声音估计词级时间戳]
|
||||||
|
C --> E
|
||||||
|
E --> F[按词的时间段汇总声音和画面特征]
|
||||||
|
D --> F
|
||||||
|
F --> G[100个样本文件 + 汇总表 + 对应关系图]
|
||||||
|
```
|
||||||
|
|
||||||
|
## 什么叫“对齐”
|
||||||
|
|
||||||
|
一段视频里,字幕、声音和画面各有自己的节奏。比如字幕中的词“happy”可能在 1.2 秒左右说出,那个时刻的视频画面里也可能出现说话人的表情变化。
|
||||||
|
|
||||||
|
“对齐”就是给不同模态的内容标上共同的时间位置,便于回答:
|
||||||
|
|
||||||
|
> 这个词出现时,声音有什么变化?同一时刻画面里有什么?
|
||||||
|
|
||||||
|
本次时间统一使用“从这条短视频开始算起的秒数”。例如 1.2 秒就是片段开始后 1.2 秒。电脑估出来的时间是模型估计值,不是人工逐帧标注的真值。
|
||||||
|
|
||||||
|
## 文字是怎么处理的
|
||||||
|
|
||||||
|
1. 直接读取题目提供的英文 transcript(字幕文本),保留原文顺序和内容,不用模型改写或纠错。
|
||||||
|
2. 按空格切成词。像 `unbelievable` 这样的词,文字模型内部可能会再拆成更小的片段。
|
||||||
|
3. 使用预训练的 BERT 英文模型,为这些更小的片段生成数字向量,再把同一个词的片段向量取平均,得到这个词的一组数字。
|
||||||
|
|
||||||
|
可以把 BERT 理解成一个读过大量英文材料的编码器:它把每个词放进一组 768 个数字里。词的向量不是情感分数,也不代表某一维就等于某种具体情绪。
|
||||||
|
|
||||||
|
## 声音是怎么处理的
|
||||||
|
|
||||||
|
声音分成两个用途不同的步骤。
|
||||||
|
|
||||||
|
### 1. 找每个词大约在什么时候说出
|
||||||
|
|
||||||
|
使用预训练的 Wav2Vec2 英语语音模型,并把题目提供的 transcript 当作已知文本。模型根据实际声音,为字幕里的字母和词寻找最可能的时间位置。这种做法叫“强制对齐”:文本固定,模型估计它在声音中的时间。
|
||||||
|
|
||||||
|
这一步不会验证字幕是不是完全正确,也不会把它当成精确的人工真值。若字幕和录音差异较大,时间估计可能不准。
|
||||||
|
|
||||||
|
### 2. 记录声音本身的变化
|
||||||
|
|
||||||
|
用 FFmpeg 把视频里的声音转成单声道、每秒 16,000 个采样点的音频,再使用 openSMILE 的 eGeMAPSv02 特征配置。它大约每 0.01 秒输出一行、每行 25 个数字,记录响度、音高、频谱形状等声音信息。
|
||||||
|
|
||||||
|
这些数值描述“声音怎么变化”,不等同于文字内容,也不是情感类别。
|
||||||
|
|
||||||
|
## 画面是怎么处理的
|
||||||
|
|
||||||
|
视频不会把每一帧都存成特征。本次每秒大约抽取 5 帧,也就是平均约每 0.2 秒取一帧。
|
||||||
|
|
||||||
|
- DeiT-Tiny 图像模型为每帧生成 192 个数字,作为通用画面特征。这组数字能表示画面信息,但单独一个数字通常没有容易读懂的名称。
|
||||||
|
- 另外,MediaPipe Face Landmarker 会尝试检测人脸,并输出 52 个脸部动作系数,例如眉毛、眼睑和嘴部的运动。这些系数描述可见的脸部动作,不是“高兴/悲伤”等情绪标签。
|
||||||
|
|
||||||
|
有 13 条视频在抽样帧中没有检测到人脸,所以这些样本没有有效的人脸动作系数;它们仍然保留了 DeiT 通用画面特征,视觉模态没有因此缺失。
|
||||||
|
|
||||||
|
## 三种特征怎样对应到词
|
||||||
|
|
||||||
|
对齐完成后,每个字幕词都有一个起止时间。例如(下面只是示意):
|
||||||
|
|
||||||
|
| 字幕词 | 估计时间 | 对应的声音 | 对应的画面 |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| happy | 1.20–1.55 秒 | 取这个时间段内的声学帧并求平均 | 取这个时间段内的视频帧特征并求平均 |
|
||||||
|
|
||||||
|
声音特征比较密,通常一个词时间段里会有多行;视频每秒只有约 5 帧,所以有些很短的词时间段里可能没有正好落入的画面帧。这种情况下,程序取时间最近的一帧,并把“使用了最近帧”的标记一并保存。
|
||||||
|
|
||||||
|
本次有 12 个词无法直接从 CTC 对齐路径中获得时间,分布在 9 条样本里。它们被放在相邻有效词之间的一个时间点上,区间长度为零,并标记 `word_alignment_valid=false`。程序没有把估计不出的词伪装成可靠的起止时间;下游对应特征通过最近的有效声音帧或画面帧取得,并保留回退标记。
|
||||||
|
|
||||||
|
## 文件里存了什么
|
||||||
|
|
||||||
|
每条视频有一个压缩的 `.npz` 文件,变长序列按原长存储,不在文件里填充到同一长度。主要数组包括:
|
||||||
|
|
||||||
|
| 内容 | 形状示例 | 含义 |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `text_features` | 词数 × 768 | 每个 transcript 词的 BERT 向量 |
|
||||||
|
| `audio_features` | 音频帧数 × 25 | 逐帧 eGeMAPS 声学特征 |
|
||||||
|
| `vision_features` | 画面帧数 × 192 | 逐帧 DeiT 通用画面特征 |
|
||||||
|
| `face_blendshape_features` | 画面帧数 × 52 | 人脸动作系数;是否有效由独立掩码表示 |
|
||||||
|
| `word_intervals_s` | 词数 × 2 | 每个词的起止时间,单位秒 |
|
||||||
|
| `audio_word_features` | 词数 × 25 | 按词的时间段汇总后的声音特征 |
|
||||||
|
| `vision_word_features` | 词数 × 192 | 按词的时间段汇总后的视频特征 |
|
||||||
|
|
||||||
|
时间戳以 32 位浮点数保存,特征值以 16 位浮点数保存以节省空间,掩码使用布尔值。掩码就是一张有效性清单:`true` 表示该位置有可用数据,`false` 表示缺失或使用了需要谨慎处理的回退结果。
|
||||||
|
|
||||||
|
## 本次处理得到的汇总
|
||||||
|
|
||||||
|
- 原始样本:100 条;标签表里的 100 个样本与视频文件一一对应,未删除或替换样本。
|
||||||
|
- 文本:1,932 个词,768 维词向量。
|
||||||
|
- 声音:77,261 个有效帧,25 维声学特征。
|
||||||
|
- 画面:3,948 帧,192 维通用视觉特征,覆盖 100/100 条样本。
|
||||||
|
- 人脸动作:87/100 条检测到有效人脸帧;其余 13 条仍有通用视觉特征。
|
||||||
|
- CTC 词级时间:1,920/1,932 个词直接对齐,12 个词带回退标记。
|
||||||
|
- 抽取错误:0 条;完整特征与报告约 10.5 MB。
|
||||||
|
|
||||||
|
审计还发现两段原始视频短于题面给出的 2.648 秒下限:`-mJ2ud6oKI8/6` 约 2.2356 秒,`-s9qJ7ATP7w/6` 约 2.4667 秒。它们按原样保留,时长差异记录在审计文件中。
|
||||||
|
|
||||||
|
## 如何复现
|
||||||
|
|
||||||
|
在 `deep_learning` 目录中运行:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd Q1
|
||||||
|
uv sync
|
||||||
|
uv run python -m q1.audit
|
||||||
|
uv run python -m q1.extract_features --output-dir outputs/q1_features
|
||||||
|
```
|
||||||
|
|
||||||
|
`uv.lock` 固定了 Python 依赖版本;运行清单记录了模型名称与修订号、参数、工具版本、GPU 信息和输出文件哈希。特征文件、汇总表、运行清单及典型样本图位于 `Q1/outputs/q1_features/`。
|
||||||
|
|
||||||
|
## 这次结果的边界
|
||||||
|
|
||||||
|
本说明讲的是“怎么从原始视频得到三模态特征,以及如何按时间组织它们”。最初的 M3/M4 版本曾把注意力摊得很平均;后来的时间码实验使 M4 找到了较稳定的时间位置,但“找到了同一时刻”仍不能保证文字、声音、画面的具体内容相互对应。情绪 Probe 也只是 100 条小样本上的检查,不能当作一般情绪识别准确率。CTC 时间、模型生成的向量和人脸动作系数都应当视为算法输出;要确认某个词和某段声音、画面是否真的对应,仍需要人工标注一些时间片段做独立核验。
|
||||||
|
|
||||||
|
## 后续的 TSFA 实验:为什么还要再对齐一次
|
||||||
|
|
||||||
|
可以把一段视频想成一条从头到尾的路。M4 根据时间先告诉我们:“第 20 个位置大约在这条路的中段。”TSFA 在中段附近再比较文字与声音、画面的内容,选择更像当前文字的片段。实验把每条视频组织成 50 个位置;“附近”在主方案中指整段视频时长的前后各 10%。如果窗口里没有可用的声音或画面位置,就取最近的有效位置,并保留这项规则。
|
||||||
|
|
||||||
|
为了知道“附近”是否真的有用,还做了两种对照:一种把同样宽的窗口随机放置,另一种让模型看完整段视频。七种方法使用相同的 100 条视频、相同的 BERT 文字特征、音频特征和 DeiT 图像特征;同一原视频的片段在同一折里一起作为训练或留出数据。总共做 5 折,每条样本都轮到一次留出评价。
|
||||||
|
|
||||||
|
结果显示,TSFA 比只用 M4 时更容易把同一位置的**文字与声音**对应起来,但画面相关的改善很小。随机窗口和看完整段视频的方案在“内容相似度”指标上甚至更高,却经常把视频前后位置弄乱。TSFA 的局部时间窗口主要帮助模型保持先后顺序。这里的“同一位置”由时间槽定义,还不是人工确认的同一事件,因此目前不能说三种模态已经准确理解了彼此。详细数字与图见 [RESULTS.md](RESULTS.md) 的 TSFA 小节。
|
||||||
@@ -0,0 +1,652 @@
|
|||||||
|
# 复杂场景下多模态情感预测:数学建模与算法设计
|
||||||
|
|
||||||
|
> 2026 年中国研究生数学建模竞赛 E 题
|
||||||
|
> **复杂场景下多模态情感预测的数学建模与算法设计**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 项目概述
|
||||||
|
|
||||||
|
本项目基于 CMU-MOSEI 英文多模态情感数据,研究复杂场景下文本(Text)、语音(Audio)和视觉(Vision)三种模态的联合情感预测。
|
||||||
|
|
||||||
|
赛题主要包含三个逐层递进的问题:
|
||||||
|
|
||||||
|
1. **Q1:多模态情感特征提取与时序对齐**
|
||||||
|
2. **Q2:模态局部缺失条件下的鲁棒情感预测**
|
||||||
|
3. **Q3:可解释性多模态情感预测**
|
||||||
|
|
||||||
|
整体逻辑可以概括为:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{
|
||||||
|
\text{原始视频}
|
||||||
|
\rightarrow
|
||||||
|
\text{多模态特征}
|
||||||
|
\rightarrow
|
||||||
|
\text{跨模态时序对齐}
|
||||||
|
\rightarrow
|
||||||
|
\text{缺失感知鲁棒融合}
|
||||||
|
\rightarrow
|
||||||
|
\text{情感预测}
|
||||||
|
\rightarrow
|
||||||
|
\text{预测解释}
|
||||||
|
}
|
||||||
|
\]
|
||||||
|
|
||||||
|
本项目现阶段不预先指定某一种方案为最终方案,而采用:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{\text{多方案全部实现 + 统一实验协议 + 系统对比}}
|
||||||
|
\]
|
||||||
|
|
||||||
|
最终根据:
|
||||||
|
|
||||||
|
- 基础预测性能;
|
||||||
|
- 缺失情况下的鲁棒性;
|
||||||
|
- 解释可信度;
|
||||||
|
- 模型复杂度;
|
||||||
|
- 可复现性;
|
||||||
|
|
||||||
|
共同确定最终建模方案。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 2. 赛题核心任务
|
||||||
|
|
||||||
|
## 2.1 Q1:多模态特征提取与时序对齐
|
||||||
|
|
||||||
|
从原始视频中分别提取:
|
||||||
|
|
||||||
|
- 文本语义特征;
|
||||||
|
- 语音情感特征;
|
||||||
|
- 视觉情感特征;
|
||||||
|
|
||||||
|
并解决三种模态:
|
||||||
|
|
||||||
|
- 采样频率不同;
|
||||||
|
- 序列长度不同;
|
||||||
|
- 时间尺度不同;
|
||||||
|
|
||||||
|
造成的时序不一致问题。
|
||||||
|
|
||||||
|
Q1 不以最终情感分类为主要目标,而应产生:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{
|
||||||
|
\text{结构化、可追溯、可核验的多模态时序特征}
|
||||||
|
}
|
||||||
|
\]
|
||||||
|
|
||||||
|
同时记录:
|
||||||
|
|
||||||
|
- 原始视频 ID;
|
||||||
|
- 模态;
|
||||||
|
- 时间位置;
|
||||||
|
- 有效长度;
|
||||||
|
- padding;
|
||||||
|
- 原始视频时间映射;
|
||||||
|
- 特征维度;
|
||||||
|
- 对齐关系。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2.2 Q2:局部模态缺失下的鲁棒情感预测
|
||||||
|
|
||||||
|
题目中的“模态缺失”不是简单的:
|
||||||
|
|
||||||
|
> 整个 Audio / Vision / Text 完全不存在。
|
||||||
|
|
||||||
|
而是:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{\text{一个或多个模态中出现随机的连续局部缺失区间}}
|
||||||
|
\]
|
||||||
|
|
||||||
|
例如:
|
||||||
|
|
||||||
|
\[
|
||||||
|
X^A =
|
||||||
|
[a_1,a_2,a_3,
|
||||||
|
\underbrace{0,0,0,0}_{\text{局部缺失}},
|
||||||
|
a_8,\ldots].
|
||||||
|
\]
|
||||||
|
|
||||||
|
模型需要同时预测:
|
||||||
|
|
||||||
|
### 情感极性
|
||||||
|
|
||||||
|
\[
|
||||||
|
\hat c \in
|
||||||
|
\{
|
||||||
|
\text{Negative},
|
||||||
|
\text{Neutral},
|
||||||
|
\text{Positive}
|
||||||
|
\}
|
||||||
|
\]
|
||||||
|
|
||||||
|
### 连续情感强度
|
||||||
|
|
||||||
|
\[
|
||||||
|
\hat y\in[-3,3]
|
||||||
|
\]
|
||||||
|
|
||||||
|
同时研究:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{
|
||||||
|
\text{缺失模态类型}
|
||||||
|
+
|
||||||
|
\text{缺失位置}
|
||||||
|
+
|
||||||
|
\text{缺失长度}
|
||||||
|
}
|
||||||
|
\]
|
||||||
|
|
||||||
|
对预测性能的影响。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2.3 Q3:可解释性情感预测
|
||||||
|
|
||||||
|
在三模态信息完整条件下,不仅预测:
|
||||||
|
|
||||||
|
\[
|
||||||
|
(\hat c,\hat y)
|
||||||
|
\]
|
||||||
|
|
||||||
|
还需要解释:
|
||||||
|
|
||||||
|
1. 主要参考了哪种模态;
|
||||||
|
2. 不同模态的作用程度;
|
||||||
|
3. 哪些局部文本、音频和视觉片段最重要;
|
||||||
|
4. 关键证据对应原始视频的什么时间位置。
|
||||||
|
|
||||||
|
因此 Q3 的核心不是简单画 Attention,而是:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{
|
||||||
|
\text{定位证据}
|
||||||
|
+
|
||||||
|
\text{量化贡献}
|
||||||
|
+
|
||||||
|
\text{验证解释}
|
||||||
|
+
|
||||||
|
\text{回溯原始素材}
|
||||||
|
}
|
||||||
|
\]
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 3. 数据集说明
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3.1 附件 1:100 条原始视频
|
||||||
|
|
||||||
|
附件 1 包含从 CMU-MOSEI 中筛选的:
|
||||||
|
|
||||||
|
\[
|
||||||
|
100
|
||||||
|
\]
|
||||||
|
|
||||||
|
条英文视频样本。
|
||||||
|
|
||||||
|
共有:
|
||||||
|
|
||||||
|
\[
|
||||||
|
37
|
||||||
|
\]
|
||||||
|
|
||||||
|
个 `video_id` 子文件夹。
|
||||||
|
|
||||||
|
样本由:
|
||||||
|
|
||||||
|
```text
|
||||||
|
video_id + clip_id
|
||||||
|
```
|
||||||
|
|
||||||
|
共同唯一确定。
|
||||||
|
|
||||||
|
视频长度范围:
|
||||||
|
|
||||||
|
\[
|
||||||
|
2.648s \sim 34.567s
|
||||||
|
\]
|
||||||
|
|
||||||
|
配套文件:
|
||||||
|
|
||||||
|
```text
|
||||||
|
label-100.xlsx
|
||||||
|
```
|
||||||
|
|
||||||
|
包含字段:
|
||||||
|
|
||||||
|
| 字段 | 含义 |
|
||||||
|
|---|---|
|
||||||
|
| video_id | 原视频编号 |
|
||||||
|
| clip_id | 视频片段编号 |
|
||||||
|
| text | 英文转写文本 |
|
||||||
|
| label | 连续情感强度 |
|
||||||
|
| annotation | Negative / Neutral / Positive |
|
||||||
|
|
||||||
|
连续情感标签:
|
||||||
|
|
||||||
|
\[
|
||||||
|
y\in[-3,3]
|
||||||
|
\]
|
||||||
|
|
||||||
|
定义:
|
||||||
|
|
||||||
|
\[
|
||||||
|
[-3,0)\Rightarrow Negative
|
||||||
|
\]
|
||||||
|
|
||||||
|
\[
|
||||||
|
0\Rightarrow Neutral
|
||||||
|
\]
|
||||||
|
|
||||||
|
\[
|
||||||
|
(0,3]\Rightarrow Positive
|
||||||
|
\]
|
||||||
|
|
||||||
|
注意:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{0\text{ 只属于 Neutral}}
|
||||||
|
\]
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 4. 附件 2:标准化多模态特征
|
||||||
|
|
||||||
|
附件 2 包含两套特征:
|
||||||
|
|
||||||
|
```text
|
||||||
|
aligned_50.pkl
|
||||||
|
unaligned_50.pkl
|
||||||
|
```
|
||||||
|
|
||||||
|
每套包含约:
|
||||||
|
|
||||||
|
\[
|
||||||
|
4850
|
||||||
|
\]
|
||||||
|
|
||||||
|
条有效样本,并按照:
|
||||||
|
|
||||||
|
```python
|
||||||
|
train
|
||||||
|
valid
|
||||||
|
test
|
||||||
|
```
|
||||||
|
|
||||||
|
划分。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4.1 Pickle 数据组织方式
|
||||||
|
|
||||||
|
整体结构为:
|
||||||
|
|
||||||
|
```text
|
||||||
|
data
|
||||||
|
├── train
|
||||||
|
│ ├── id
|
||||||
|
│ ├── raw_text
|
||||||
|
│ ├── text
|
||||||
|
│ ├── text_bert
|
||||||
|
│ ├── audio
|
||||||
|
│ ├── vision
|
||||||
|
│ ├── annotations
|
||||||
|
│ ├── classification_labels
|
||||||
|
│ ├── regression_labels
|
||||||
|
│ └── ...
|
||||||
|
├── valid
|
||||||
|
└── test
|
||||||
|
```
|
||||||
|
|
||||||
|
正确读取方式:
|
||||||
|
|
||||||
|
```python
|
||||||
|
data["train"]["audio"][j]
|
||||||
|
```
|
||||||
|
|
||||||
|
而不是:
|
||||||
|
|
||||||
|
```python
|
||||||
|
data["train"][j]["audio"]
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 5. aligned / unaligned 数据格式
|
||||||
|
|
||||||
|
## 5.1 aligned_50.pkl
|
||||||
|
|
||||||
|
\[
|
||||||
|
X^T\in\mathbb R^{N\times50\times768}
|
||||||
|
\]
|
||||||
|
|
||||||
|
\[
|
||||||
|
X^A\in\mathbb R^{N\times50\times74}
|
||||||
|
\]
|
||||||
|
|
||||||
|
\[
|
||||||
|
X^V\in\mathbb R^{N\times50\times35}
|
||||||
|
\]
|
||||||
|
|
||||||
|
三个模态已经被组织为:
|
||||||
|
|
||||||
|
\[
|
||||||
|
50
|
||||||
|
\]
|
||||||
|
|
||||||
|
个相互对应的位置。
|
||||||
|
|
||||||
|
可以理解为:
|
||||||
|
|
||||||
|
\[
|
||||||
|
T_i\leftrightarrow A_i\leftrightarrow V_i.
|
||||||
|
\]
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5.2 unaligned_50.pkl
|
||||||
|
|
||||||
|
文本:
|
||||||
|
|
||||||
|
\[
|
||||||
|
X^T\in\mathbb R^{N\times50\times768}
|
||||||
|
\]
|
||||||
|
|
||||||
|
音频:
|
||||||
|
|
||||||
|
\[
|
||||||
|
X^A\in\mathbb R^{N\times500\times74}
|
||||||
|
\]
|
||||||
|
|
||||||
|
视觉:
|
||||||
|
|
||||||
|
\[
|
||||||
|
X^V\in\mathbb R^{N\times500\times35}
|
||||||
|
\]
|
||||||
|
|
||||||
|
同时提供:
|
||||||
|
|
||||||
|
```text
|
||||||
|
audio_lengths
|
||||||
|
vision_lengths
|
||||||
|
```
|
||||||
|
|
||||||
|
记录实际有效长度。
|
||||||
|
|
||||||
|
因此 unaligned 数据保留了:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{
|
||||||
|
\text{文本较粗语义轴}
|
||||||
|
+
|
||||||
|
\text{音频细粒度时序}
|
||||||
|
+
|
||||||
|
\text{视觉细粒度时序}
|
||||||
|
}
|
||||||
|
\]
|
||||||
|
|
||||||
|
这为学习式 Cross-Attention 对齐提供了空间。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 6. 附件 3:模态局部缺失测试集
|
||||||
|
|
||||||
|
附件 3:
|
||||||
|
|
||||||
|
- 无标签;
|
||||||
|
- 已处理多模态特征;
|
||||||
|
- 包含随机局部模态缺失;
|
||||||
|
- 缺失区间表现为连续位置全部置零。
|
||||||
|
|
||||||
|
用于 Q2 最终预测。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 7. 附件 4:可解释性专项测试集
|
||||||
|
|
||||||
|
附件 4:
|
||||||
|
|
||||||
|
- 无标签;
|
||||||
|
- 三模态完整;
|
||||||
|
- 同时提供原始真实场景视频;
|
||||||
|
- 特征组织形式与附件 2 一致。
|
||||||
|
|
||||||
|
用于:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{
|
||||||
|
\text{预测}
|
||||||
|
+
|
||||||
|
\text{模态贡献}
|
||||||
|
+
|
||||||
|
\text{关键证据定位}
|
||||||
|
}
|
||||||
|
\]
|
||||||
|
|
||||||
|
可以根据 Q1 保存的时间映射将关键位置重新定位到:
|
||||||
|
|
||||||
|
- 原始文本;
|
||||||
|
- 音频时间段;
|
||||||
|
- 视频关键帧。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 8. 统一数学符号
|
||||||
|
|
||||||
|
一个样本记为:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\mathcal X_i=
|
||||||
|
(X_i^T,X_i^A,X_i^V)
|
||||||
|
\]
|
||||||
|
|
||||||
|
其中:
|
||||||
|
|
||||||
|
- \(T\):Text
|
||||||
|
- \(A\):Audio
|
||||||
|
- \(V\):Vision
|
||||||
|
|
||||||
|
原始模态序列:
|
||||||
|
|
||||||
|
\[
|
||||||
|
X^m=
|
||||||
|
[x_1^m,\ldots,x_{L_m}^m],
|
||||||
|
\qquad
|
||||||
|
m\in\{T,A,V\}.
|
||||||
|
\]
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8.1 基础符号
|
||||||
|
|
||||||
|
| 符号 | 含义 |
|
||||||
|
|---|---|
|
||||||
|
| \(x_j^m\) | 模态 \(m\) 原始第 \(j\) 个位置特征 |
|
||||||
|
| \(\tilde x_i^m\) | 软对齐到文本位置 \(i\) 后的特征 |
|
||||||
|
| \(h_i^m\) | 模态编码器生成的隐藏表示 |
|
||||||
|
| \(M_j^m\) | 原始位置是否有效 |
|
||||||
|
| \(A_{ij}^m\) | Cross-Attention 对齐权重 |
|
||||||
|
| \(R_i^m\) | 对齐后的局部模态可靠度 |
|
||||||
|
| \(\alpha_i^m\) | 融合时模态权重 |
|
||||||
|
| \(z_i\) | 第 \(i\) 个位置的多模态融合表示 |
|
||||||
|
| \(\beta_i\) | 第 \(i\) 个位置对整段情绪的时间重要性 |
|
||||||
|
| \(g\) | 整条视频最终表示 |
|
||||||
|
| \(\hat y\) | 情感强度预测 |
|
||||||
|
| \(\hat c\) | 情感极性预测 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Q1
|
||||||
|
@Q1.md
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 可能使用的开源库、工具与模型
|
||||||
|
|
||||||
|
| 类别 | 工具 / 模型 | 用途 | 优先级 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| 深度学习 | PyTorch | 网络、训练、Cross-Attention | 必须 |
|
||||||
|
| GPU | CUDA / cuDNN | GPU 加速 | 必须 |
|
||||||
|
| 数值计算 | NumPy | 数值运算 | 必须 |
|
||||||
|
| 数据 | Pandas | CSV / XLSX / 表格 | 必须 |
|
||||||
|
| 科学计算 | SciPy | Pearson / 统计分析 | 必须 |
|
||||||
|
| ML | scikit-learn | F1、Accuracy、MAE 等 | 必须 |
|
||||||
|
| 配置 | PyYAML | 实验配置 | 推荐 |
|
||||||
|
| 配置 | Hydra / OmegaConf | 多实验管理 | 推荐 |
|
||||||
|
| 日志 | TensorBoard | loss / metric | 推荐 |
|
||||||
|
| 日志 | MLflow | 系统实验管理 | 可选 |
|
||||||
|
| 调参 | Optuna | 超参数搜索 | 推荐 |
|
||||||
|
| 视频 | FFmpeg | 拆音频、转码 | 必须 |
|
||||||
|
| 视频 | OpenCV | 抽帧与时间映射 | 必须 |
|
||||||
|
| 文本 | Hugging Face Transformers | BERT 等 | 高 |
|
||||||
|
| 文本 | BERT-base-uncased | 768 维表示 | 高 |
|
||||||
|
| ASR/时间戳 | WhisperX | 词级时间定位 | 推荐 |
|
||||||
|
| 强制对齐 | Montreal Forced Aligner | transcript-audio alignment | 推荐 |
|
||||||
|
| 音频 | openSMILE | 声学/韵律特征 | 高 |
|
||||||
|
| 音频 | librosa | 基础信号处理 | 推荐 |
|
||||||
|
| 音频 | torchaudio | PyTorch 音频 | 推荐 |
|
||||||
|
| 音频 PLM | wav2vec 2.0 | 深层语音表示 | 可选 |
|
||||||
|
| 音频 PLM | WavLM | 深层语音表示 | 可选 |
|
||||||
|
| 音频 PLM | HuBERT | 深层语音表示 | 可选 |
|
||||||
|
| 视觉 | OpenFace 2.0 | AU / gaze / pose | 高 |
|
||||||
|
| 视觉 | MediaPipe | landmark / blendshape | 可选 |
|
||||||
|
| 视觉 | torchvision | 图像处理 | 推荐 |
|
||||||
|
| 多模态 | CMU-MultimodalSDK | MOSEI 数据参考 | 高 |
|
||||||
|
| 多模态 | MultiBench | baseline / multimodal benchmark | 高 |
|
||||||
|
| 解释 | Captum | IG / Occlusion / Ablation | 高 |
|
||||||
|
| 解释 | SHAP | Shapley 解释 | 可选 |
|
||||||
|
| 可视化 | Matplotlib | 曲线 / 热图 | 必须 |
|
||||||
|
| 可视化 | Plotly | 交互可视化 | 可选 |
|
||||||
|
| 统计 | statsmodels | 回归 / 显著性分析 | 推荐 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 模型候选汇总
|
||||||
|
|
||||||
|
| 模块 | 候选 |
|
||||||
|
|---|---|
|
||||||
|
| Text Encoder | BERT |
|
||||||
|
| Audio Feature | openSMILE |
|
||||||
|
| Audio Encoder | BiGRU / Transformer |
|
||||||
|
| Audio PLM | Wav2Vec2 / WavLM |
|
||||||
|
| Vision Feature | OpenFace |
|
||||||
|
| Vision Encoder | BiGRU / Transformer |
|
||||||
|
| Alignment | Hard Alignment |
|
||||||
|
| Alignment | Fixed Window |
|
||||||
|
| Alignment | Cross-Attention |
|
||||||
|
| Fusion | Concatenation |
|
||||||
|
| Fusion | TFN |
|
||||||
|
| Fusion | LMF |
|
||||||
|
| Fusion | MAG |
|
||||||
|
| Fusion | MulT |
|
||||||
|
| Representation | MISA |
|
||||||
|
| Missing Model | MMIN |
|
||||||
|
| Missing Model | CMAD |
|
||||||
|
| Missing Model | P-RMF |
|
||||||
|
| Experts | EMOE |
|
||||||
|
| Diffusion | HyperEF |
|
||||||
|
| Factorization | FUSE-Net |
|
||||||
|
| Distribution Alignment | CaReFlow |
|
||||||
|
| Robust Representation | CmIR |
|
||||||
|
| Attribution | Integrated Gradients |
|
||||||
|
| Attribution | Occlusion |
|
||||||
|
| Attribution | Feature Ablation |
|
||||||
|
| Attribution | SHAP |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 参考文献
|
||||||
|
|
||||||
|
> 下表优先使用论文官方会议/期刊入口。
|
||||||
|
> README 中记录 DOI、ACL Anthology ID、arXiv ID 或官方论文标题,避免引用二手博客作为正式论文来源。
|
||||||
|
|
||||||
|
| # | 文献 | 会议/期刊 | 年份 | 与本项目关系 | 官方标识 |
|
||||||
|
|---|---|---|---:|---|---|
|
||||||
|
| 1 | Zhang et al., *Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects* | Expert Systems with Applications | 2024 | 多模态情感综述 | DOI: 10.1016/j.eswa.2023.121692 |
|
||||||
|
| 2 | Bagher Zadeh et al., *Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph* | ACL | 2018 | CMU-MOSEI | ACL: P18-1208 |
|
||||||
|
| 3 | Vaswani et al., *Attention Is All You Need* | NeurIPS | 2017 | Transformer / Attention | NeurIPS 2017 |
|
||||||
|
| 4 | Devlin et al., *BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding* | NAACL | 2019 | 文本特征 | ACL: N19-1423 |
|
||||||
|
| 5 | Eyben et al., *openSMILE: The Munich Versatile and Fast Open-Source Audio Feature Extractor* | ACM Multimedia | 2010 | 音频特征 | ACM MM 2010 |
|
||||||
|
| 6 | Baltrušaitis et al., *OpenFace 2.0: Facial Behavior Analysis Toolkit* | FG | 2018 | 视觉特征 | DOI: 10.1109/FG.2018.00019 |
|
||||||
|
| 7 | Bain et al., *WhisperX: Time-Accurate Speech Transcription of Long-Form Audio* | arXiv | 2023 | 词级时间戳 | arXiv:2303.00747 |
|
||||||
|
| 8 | Tsai et al., *Multimodal Transformer for Unaligned Multimodal Language Sequences* | ACL | 2019 | unaligned + Cross-Attention | ACL: P19-1656 |
|
||||||
|
| 9 | Rahman et al., *Integrating Multimodal Information in Large Pretrained Transformers* | ACL | 2020 | MAG-BERT / 文本中心融合 | DOI:10.18653/v1/2020.acl-main.214 |
|
||||||
|
| 10 | Zadeh et al., *Tensor Fusion Network for Multimodal Sentiment Analysis* | EMNLP | 2017 | TFN baseline | ACL: D17-1115 |
|
||||||
|
| 11 | Liu et al., *Efficient Low-rank Multimodal Fusion With Modality-Specific Factors* | ACL | 2018 | LMF baseline | ACL: P18-1209 |
|
||||||
|
| 12 | Hazarika et al., *MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis* | ACM MM | 2020 | shared / private representation | DOI:10.1145/3394171.3413678 |
|
||||||
|
| 13 | Neverova et al., *ModDrop: Adaptive Multi-Modal Gesture Recognition* | TPAMI | 2016 | 模态随机丢失 | DOI:10.1109/TPAMI.2015.2461544 |
|
||||||
|
| 14 | Pham et al., *Found in Translation: Learning Robust Joint Representations by Cyclic Translations between Modalities* | AAAI | 2019 | 跨模态重构 | DOI:10.1609/aaai.v33i01.33016892 |
|
||||||
|
| 15 | Zhao et al., *Missing Modality Imagination Network for Emotion Recognition with Uncertain Missing Modalities* | ACL-IJCNLP | 2021 | MMIN | ACL:2021.acl-long.203 |
|
||||||
|
| 16 | 王楠、王淇、欧阳丹彤,《基于知识蒸馏与动态调整机制的多模态情感分析模型》 | 计算机学报 | 2025 | AUMDF / 动态权重 / 缺失 | 48(8):1923–1942 |
|
||||||
|
| 17 | Fang et al., *EMOE: Modality-Specific Enhanced Dynamic Emotion Experts* | CVPR | 2025 | Mixture of Experts | CVPR 2025 |
|
||||||
|
| 18 | Zhuang et al., *CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing Modalities* | ICCV | 2025 | 缺失模态蒸馏 | ICCV 2025 |
|
||||||
|
| 19 | Zhu et al., *Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete Data* | ACL | 2025 | P-RMF / 不确定性 | DOI:10.18653/v1/2025.acl-long.1075 |
|
||||||
|
| 20 | Qiu et al., *Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition* | CVPR | 2026 | Diffusion / uncertainty | CVPR 2026 |
|
||||||
|
| 21 | Yang & Li, *Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment Analysis* | CVPR | 2026 | shared/specific/noise + reconstruction | CVPR 2026 |
|
||||||
|
| 22 | Mai & Han, *Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference Perspective* | ACL | 2026 | CmIR / causal invariance | DOI:10.18653/v1/2026.acl-long.2119 |
|
||||||
|
| 23 | Mai & Han, *CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion* | CVPR | 2026 | 模态分布对齐 | arXiv:2602.19140 |
|
||||||
|
| 24 | Wan et al., *Locate and Explain: Joint Multimodal Emotion Cause Extraction and Summarization in Conversation* | ACL | 2026 | 关键证据定位 | DOI:10.18653/v1/2026.acl-long.2012 |
|
||||||
|
| 25 | Sundararajan et al., *Axiomatic Attribution for Deep Networks* | ICML | 2017 | Integrated Gradients | PMLR 70 |
|
||||||
|
| 26 | Jain & Wallace, *Attention is not Explanation* | NAACL | 2019 | Attention 解释局限 | DOI:10.18653/v1/N19-1357 |
|
||||||
|
| 27 | Lundberg & Lee, *A Unified Approach to Interpreting Model Predictions* | NeurIPS | 2017 | SHAP | NeurIPS 2017 |
|
||||||
|
| 28 | Adebayo et al., *Sanity Checks for Saliency Maps* | NeurIPS | 2018 | 解释可信性验证 | NeurIPS 2018 |
|
||||||
|
| 29 | Zeiler & Fergus, *Visualizing and Understanding Convolutional Networks* | ECCV | 2014 | Occlusion 思想 | DOI:10.1007/978-3-319-10590-1_53 |
|
||||||
|
| 30 | Wachter et al., *Counterfactual Explanations without Opening the Black Box* | arXiv / HILDA | 2017 | 反事实解释 | arXiv:1711.00399 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# 最终原则
|
||||||
|
|
||||||
|
本项目当前不采用:
|
||||||
|
|
||||||
|
> “先选一个看起来高级的模型,然后证明它最好。”
|
||||||
|
|
||||||
|
而采用:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{
|
||||||
|
\text{提出多个机制假设}
|
||||||
|
\rightarrow
|
||||||
|
\text{设计公平实验}
|
||||||
|
\rightarrow
|
||||||
|
\text{观察数据}
|
||||||
|
\rightarrow
|
||||||
|
\text{解释规律}
|
||||||
|
\rightarrow
|
||||||
|
\text{确定最终模型}
|
||||||
|
}
|
||||||
|
\]
|
||||||
|
|
||||||
|
重点不只是得到最高指标,还要回答:
|
||||||
|
|
||||||
|
1. 为什么某种对齐方式更好?
|
||||||
|
2. 哪种模态最怕局部缺失?
|
||||||
|
3. 缺失的位置是否比缺失比例更重要?
|
||||||
|
4. Cross-Attention 能否帮助判断缺失信息的重要程度?
|
||||||
|
5. 动态可靠度是否真的提高鲁棒性?
|
||||||
|
6. 重构和“降低信任”哪种策略更有效?
|
||||||
|
7. Attention 权重是否真的对应模型决策依据?
|
||||||
|
8. 哪种解释方法最能经受反事实验证?
|
||||||
|
|
||||||
|
最终目标是形成:
|
||||||
|
|
||||||
|
\[
|
||||||
|
\boxed{
|
||||||
|
\text{有性能}
|
||||||
|
+
|
||||||
|
\text{有机制}
|
||||||
|
+
|
||||||
|
\text{有数学分析}
|
||||||
|
+
|
||||||
|
\text{有可解释性}
|
||||||
|
}
|
||||||
|
\]
|
||||||
|
|
||||||
|
的完整多模态情感预测建模方案。
|
||||||
Binary file not shown.
|
After Width: | Height: | Size: 25 KiB |
@@ -0,0 +1,44 @@
|
|||||||
|
# 问题一模型对比
|
||||||
|
|
||||||
|
本目录实现问题一 PDF 第 4.8.3 节的 B0–B4 单因素对照。代码从附件一的 100 条原始视频和标签表开始,先生成统一的文本、语音、视觉特征,再在 0.1 秒主时间网格上进行五折 `video_id` 留组评估。
|
||||||
|
|
||||||
|
## 运行
|
||||||
|
|
||||||
|
在仓库根目录执行:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd math
|
||||||
|
uv sync
|
||||||
|
uv run python compare_models.py
|
||||||
|
```
|
||||||
|
|
||||||
|
默认使用 CUDA(如果可用),需要可访问 Hugging Face 下载 BERT 与 Wav2Vec2 权重;MediaPipe Face Landmarker 权重会保存在 `math/cache/`。可用 `--device cpu` 指定 CPU。缓存保存在 `math/cache/native/`,中断后重跑会继续使用已完成样本。只有传入 `--force-extract` 才会重算全部原始特征。
|
||||||
|
|
||||||
|
## 对照定义
|
||||||
|
|
||||||
|
- **B0**:冻结 BERT 词向量、冻结 CTC Viterbi 词边界、物理时间区间投影;训练折均值/标准差缩放。
|
||||||
|
- **B1**:B0 加训练折中位数/MAD 稳健缩放,并对面部旋转使用 SO(3) 均值、对视线代理使用单位向量均值。
|
||||||
|
- **B2**:B1 只替换文本时间投影,使用 CTC 固定转写状态图的前向–后向词占据后验。音频/视频的 PTS 不变。
|
||||||
|
- **B3**:B1 加音频 log-F0/能量和面部 jaw-open/嘴部开合率的一、二阶时间增广路径签名。缺测处会断开路径。
|
||||||
|
- **B4**:B1 加冻结 Wav2Vec2-base-960h 最后四层的时间投影表示。
|
||||||
|
|
||||||
|
每个模型以相同的五段时间池化探针评价极性和强度。分类器为固定 C=0.05 的逻辑回归,回归器为固定 alpha=25 的 Ridge;参数不按留出折调节。归一化器只使用训练折观测值。95% 差值区间以 `video_id` 为单位做 2,000 次配对 Bootstrap。
|
||||||
|
|
||||||
|
文本维度为 768(BERT 最后四层、词内子词平均);语音为 74 维(40 log-Mel、13 MFCC、13 ΔMFCC、8 韵律/谱量);视觉为 35 维(17 个 MediaPipe blendshape 代理、6 维头姿、6 维近似视线、6 维面部比例几何)。MediaPipe blendshape 分数不等于 OpenFace AU 强度,视线为基于虹膜偏移的代理量;本次实现记录了实际所用索引与定义,没有将代理量称为 OpenFace 特征。
|
||||||
|
|
||||||
|
硬边界质量由 CTC 发射分数得到,但不视为校准概率。聚合时按样本内有效词的中位分数归一化,并截断到 [0.25, 4],仅作为相对质量权重。物理覆盖率仍以实际区间交叠计算。B2 的文本后验占据率是期望词活动质量,不能与物理覆盖率互换。
|
||||||
|
|
||||||
|
没有人工词边界标注,因此结果不报告边界误差、IoU/MATE 或后验经验校准率。情感探针分数只检查当前特征视图保留标签信息的程度,不是对齐正确性的真值。
|
||||||
|
|
||||||
|
## 结果文件
|
||||||
|
|
||||||
|
完整结果在 [`results/model_comparison/`](results/model_comparison/):
|
||||||
|
|
||||||
|
- `report.md`:实验设置、主要结果和限制。
|
||||||
|
- `comparison_summary.csv`、`fold_metrics.csv`:总体和分折指标。
|
||||||
|
- `group_bootstrap_deltas.csv`:相对 B1 的视频组配对 Bootstrap 区间。
|
||||||
|
- `oof_predictions.csv`:100 条留组预测。
|
||||||
|
- `sample_alignment_summary.csv`、`modality_summary.csv`:100 行样本汇总与 300 行模态明细。
|
||||||
|
- `word_alignment_posterior.csv`:逐词硬边界、未校准对齐分数、实际相对质量权重与模型内部边界区间。
|
||||||
|
- `split_assignments.csv`、`run_manifest.json`:折号、文件哈希、模型 revision、运行环境和参数。
|
||||||
|
- `comparison.png`:OOF 指标图。
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,24 @@
|
|||||||
|
[project]
|
||||||
|
name = "math-q1-comparison"
|
||||||
|
version = "0.1.0"
|
||||||
|
requires-python = ">=3.14"
|
||||||
|
dependencies = [
|
||||||
|
"matplotlib>=3.11.2",
|
||||||
|
"mediapipe>=1.0.1",
|
||||||
|
"numpy>=2.5.2",
|
||||||
|
"opencv-python-headless>=5.0.0.93",
|
||||||
|
"openpyxl>=3.1.5",
|
||||||
|
"pillow>=12.3.0",
|
||||||
|
"scikit-learn>=1.9.1",
|
||||||
|
"scipy>=1.17.1",
|
||||||
|
"torch>=2.14.0",
|
||||||
|
"transformers>=5.17.0",
|
||||||
|
]
|
||||||
|
|
||||||
|
[tool.uv.sources]
|
||||||
|
torch = { index = "pytorch" }
|
||||||
|
|
||||||
|
[[tool.uv.index]]
|
||||||
|
name = "pytorch"
|
||||||
|
url = "https://download.pytorch.org/whl/cu130"
|
||||||
|
explicit = true
|
||||||
Binary file not shown.
|
After Width: | Height: | Size: 100 KiB |
@@ -0,0 +1,6 @@
|
|||||||
|
method,sample_count,video_group_count,oof_accuracy,oof_macro_f1,oof_mae,oof_pearson,fold_accuracy_mean,fold_accuracy_sd,fold_macro_f1_mean,fold_macro_f1_sd,fold_mae_mean,fold_mae_sd,fold_pearson_mean,fold_pearson_sd
|
||||||
|
B0,100,37,0.6,0.3612633181126332,0.5642678308486938,0.34585652345724877,0.6,0.1274754878398196,0.33701813174429807,0.09810807621270928,0.5642678308486938,0.15030090846290933,0.40788979900767275,0.06001860958324475
|
||||||
|
B1,100,37,0.61,0.387431302270012,0.5989421024918556,0.23765809617297712,0.61,0.1341640786499874,0.3493607699490052,0.11905440418044842,0.5989421024918556,0.15750913261096855,0.281166840616461,0.14435440813423514
|
||||||
|
B2,100,37,0.58,0.33061794334825517,0.588000754788518,0.2790493470006418,0.5800000000000001,0.14404860290887933,0.3128829195495862,0.08742851668319461,0.588000754788518,0.14204292727229248,0.3331122890661654,0.08879703094318196
|
||||||
|
B3,100,37,0.59,0.3587962962962963,0.5997931832820177,0.23011191320162028,0.5900000000000001,0.12942179105544785,0.33257284523004604,0.09564119982991781,0.5997931832820177,0.15746511520951317,0.2803250178248575,0.13217248970127823
|
||||||
|
B4,100,37,0.6,0.4222603610475464,0.59276591360569,0.19864510121331094,0.6,0.13693063937629155,0.3976841268630668,0.08882960164587973,0.59276591360569,0.16014796281119822,0.24000288610371706,0.14964617105964673
|
||||||
|
@@ -0,0 +1,26 @@
|
|||||||
|
method,fold,train_samples,valid_samples,train_video_groups,valid_video_groups,accuracy,macro_f1,mae,pearson,feature_dimension
|
||||||
|
B0,1,80,20,31,6,0.75,0.29411764705882354,0.4224671095609665,0.39024875749849447,4385
|
||||||
|
B1,1,80,20,31,6,0.75,0.29411764705882354,0.43092925772070884,0.3768780281745036,4385
|
||||||
|
B2,1,80,20,31,6,0.75,0.2857142857142857,0.43804959058761594,0.3386523481377315,4385
|
||||||
|
B3,1,80,20,31,6,0.75,0.29411764705882354,0.42625029012560844,0.3688530099245201,4455
|
||||||
|
B4,1,80,20,31,6,0.8,0.42745098039215684,0.41526668667793276,0.3438838954822536,8225
|
||||||
|
B0,2,80,20,30,7,0.6,0.3639846743295019,0.776566531509161,0.3958411255263523,4385
|
||||||
|
B1,2,80,20,30,7,0.6,0.35714285714285715,0.8237437169998885,0.24091732574554836,4385
|
||||||
|
B2,2,80,20,30,7,0.6,0.35714285714285715,0.7888965889811516,0.3639766310607314,4385
|
||||||
|
B3,2,80,20,30,7,0.6,0.35714285714285715,0.8207762833684683,0.256615486910397,4455
|
||||||
|
B4,2,80,20,30,7,0.6,0.3639846743295019,0.8059555269777775,0.19887966518486633,8225
|
||||||
|
B0,3,80,20,29,8,0.65,0.44334975369458124,0.6439033597707748,0.4541649022346378,4385
|
||||||
|
B1,3,80,20,29,8,0.7,0.5119047619047619,0.6599065795540809,0.3794933845977076,4385
|
||||||
|
B2,3,80,20,29,8,0.6,0.35555555555555557,0.660385686159134,0.3833892851036205,4385
|
||||||
|
B3,3,80,20,29,8,0.65,0.44334975369458124,0.6637263983488083,0.36870656729912815,4455
|
||||||
|
B4,3,80,20,29,8,0.5,0.3956043956043956,0.6777055777609349,0.23098821201949785,8225
|
||||||
|
B0,4,80,20,29,8,0.4,0.19047619047619047,0.5507291875779629,0.32319803952574416,4385
|
||||||
|
B1,4,80,20,29,8,0.4,0.19047619047619047,0.6101470191031695,0.04446251878558958,4385
|
||||||
|
B2,4,80,20,29,8,0.35,0.1728395061728395,0.5764139499515295,0.17960866044359325,4385
|
||||||
|
B3,4,80,20,29,8,0.4,0.19047619047619047,0.6146079197525978,0.0589277387406542,4455
|
||||||
|
B4,4,80,20,29,8,0.45,0.2792022792022792,0.607093845307827,0.018750134966085342,8225
|
||||||
|
B0,5,80,20,29,8,0.6,0.39316239316239315,0.427672965824604,0.4759961702531348,4385
|
||||||
|
B1,5,80,20,29,8,0.6,0.39316239316239315,0.46998393908143044,0.3640829457789559,4385
|
||||||
|
B2,5,80,20,29,8,0.6,0.39316239316239315,0.4762579582631588,0.3999345205851506,4385
|
||||||
|
B3,5,80,20,29,8,0.55,0.37777777777777777,0.4736050248146057,0.34852228624958836,4455
|
||||||
|
B4,5,80,20,29,8,0.65,0.5221783047870004,0.45780793130397796,0.4075125228658822,8225
|
||||||
|
@@ -0,0 +1,17 @@
|
|||||||
|
comparison,metric,reference_oof,method_oof,delta_oof,group_bootstrap_ci95_low,group_bootstrap_ci95_high,bootstrap_repeats
|
||||||
|
B0-B1,accuracy,0.61,0.6,-0.010000000000000009,-0.029710396039603987,0.0,2000
|
||||||
|
B0-B1,macro_f1,0.387431302270012,0.3612633181126332,-0.02616798415737881,-0.08327057744765032,0.0,2000
|
||||||
|
B0-B1,mae,0.5989421024918556,0.5642678308486938,-0.034674271643161836,-0.06304207975806683,-0.012443204964559182,2000
|
||||||
|
B0-B1,pearson,0.23765809617297712,0.34585652345724877,0.10819842728427165,0.04340079372240924,0.19005024141624596,2000
|
||||||
|
B2-B1,accuracy,0.61,0.58,-0.030000000000000027,-0.0648204607046071,0.0,2000
|
||||||
|
B2-B1,macro_f1,0.387431302270012,0.33061794334825517,-0.05681335892175682,-0.10940517350156256,-0.0009457159449812041,2000
|
||||||
|
B2-B1,mae,0.5989421024918556,0.588000754788518,-0.010941347703337656,-0.027305471624463586,0.004017100735475708,2000
|
||||||
|
B2-B1,pearson,0.23765809617297712,0.2790493470006418,0.041391250827664705,-0.009244736549056344,0.10131359014731482,2000
|
||||||
|
B3-B1,accuracy,0.61,0.59,-0.020000000000000018,-0.05154639175257736,0.0,2000
|
||||||
|
B3-B1,macro_f1,0.387431302270012,0.3587962962962963,-0.02863500597371571,-0.07212177947137302,0.0,2000
|
||||||
|
B3-B1,mae,0.5989421024918556,0.5997931832820177,0.000851080790162051,-0.004701075716562739,0.006246566925533936,2000
|
||||||
|
B3-B1,pearson,0.23765809617297712,0.23011191320162028,-0.007546182971356841,-0.022226358336772584,0.005402452990031458,2000
|
||||||
|
B4-B1,accuracy,0.61,0.6,-0.010000000000000009,-0.08080808080808077,0.06196073008849559,2000
|
||||||
|
B4-B1,macro_f1,0.387431302270012,0.4222603610475464,0.034829058777534394,-0.051408708516982864,0.13655330256210496,2000
|
||||||
|
B4-B1,mae,0.5989421024918556,0.59276591360569,-0.006176188886165668,-0.03045045264816862,0.015397113913740889,2000
|
||||||
|
B4-B1,pearson,0.23765809617297712,0.19864510121331094,-0.039012994959666175,-0.1082683426423725,0.03034022099301614,2000
|
||||||
|
@@ -0,0 +1,301 @@
|
|||||||
|
sample_id,video_id,clip_id,modality,source_duration_s,observed_duration_s_mean_dimension,native_length,native_dimension,main_grid_length,grid_step_s,grid_dimension,mean_grid_coverage,status,source_video_sha256
|
||||||
|
-3g5yACwYnA/13,-3g5yACwYnA,13,text,5.5139970779418945,2.946523889899254,15,768,56,0.1,768,0.7015532851219177,ok,aeb47627f59dce3d0a85a44ef35e4a3bc18211498e98c8c11cf60269646df24f
|
||||||
|
-3g5yACwYnA/13,-3g5yACwYnA,13,audio,5.5139970779418945,5.445240411887298,540,74,56,0.1,74,0.9882608652114868,ok,aeb47627f59dce3d0a85a44ef35e4a3bc18211498e98c8c11cf60269646df24f
|
||||||
|
-3g5yACwYnA/13,-3g5yACwYnA,13,vision,5.5139970779418945,5.5139970779418945,43,35,56,0.1,35,1.0,ok,aeb47627f59dce3d0a85a44ef35e4a3bc18211498e98c8c11cf60269646df24f
|
||||||
|
-3g5yACwYnA/3,-3g5yACwYnA,3,text,14.388997077941895,7.041063531115653,29,768,144,0.1,768,0.6343300342559814,ok,eff8cfefba2425155f2b72656829b34b23be8bfe1dcbece384154f56818aa26d
|
||||||
|
-3g5yACwYnA/3,-3g5yACwYnA,3,audio,14.388997077941895,14.260037878559858,1433,74,144,0.1,74,0.9912122488021851,ok,eff8cfefba2425155f2b72656829b34b23be8bfe1dcbece384154f56818aa26d
|
||||||
|
-3g5yACwYnA/3,-3g5yACwYnA,3,vision,14.388997077941895,14.388997077941895,103,35,144,0.1,35,1.0,ok,eff8cfefba2425155f2b72656829b34b23be8bfe1dcbece384154f56818aa26d
|
||||||
|
-3g5yACwYnA/2,-3g5yACwYnA,2,text,9.394009590148926,5.1430321040563305,14,768,94,0.1,768,0.7563282251358032,ok,0619c01f137d017c15c058176f18a575919189c7e81ebd2bfb993afdb86f3de7
|
||||||
|
-3g5yACwYnA/2,-3g5yACwYnA,2,audio,9.394009590148926,9.331671435768538,924,74,94,0.1,74,0.9936367273330688,ok,0619c01f137d017c15c058176f18a575919189c7e81ebd2bfb993afdb86f3de7
|
||||||
|
-3g5yACwYnA/2,-3g5yACwYnA,2,vision,9.394009590148926,9.394009590148926,78,35,94,0.1,35,1.0,ok,0619c01f137d017c15c058176f18a575919189c7e81ebd2bfb993afdb86f3de7
|
||||||
|
-3g5yACwYnA/9,-3g5yACwYnA,9,text,8.816991806030273,5.6415157541632714,21,768,89,0.1,768,0.6637077927589417,ok,b962573b12ed1f06d5533f6427c80dce2bae3a99dd332f6d3fcb32c579f3f016
|
||||||
|
-3g5yACwYnA/9,-3g5yACwYnA,9,audio,8.816991806030273,8.735654204761659,870,74,89,0.1,74,0.9912108778953552,ok,b962573b12ed1f06d5533f6427c80dce2bae3a99dd332f6d3fcb32c579f3f016
|
||||||
|
-3g5yACwYnA/9,-3g5yACwYnA,9,vision,8.816991806030273,8.816991806030273,75,35,89,0.1,35,1.0,ok,b962573b12ed1f06d5533f6427c80dce2bae3a99dd332f6d3fcb32c579f3f016
|
||||||
|
-3nNcZdcdvU/5,-3nNcZdcdvU,5,text,7.867969036102295,5.002839090395719,18,768,79,0.1,768,0.7357116341590881,ok,7ea53ad502b77be7d2ee4494dc23a9478c897792203cf887d546e83d584c1728
|
||||||
|
-3nNcZdcdvU/5,-3nNcZdcdvU,5,audio,7.867969036102295,7.843361772234375,779,74,79,0.1,74,0.9971166849136353,ok,7ea53ad502b77be7d2ee4494dc23a9478c897792203cf887d546e83d584c1728
|
||||||
|
-3nNcZdcdvU/5,-3nNcZdcdvU,5,vision,7.867969036102295,7.816666602730405,67,35,79,0.1,35,0.9904456734657288,ok,7ea53ad502b77be7d2ee4494dc23a9478c897792203cf887d546e83d584c1728
|
||||||
|
-HwX2H8Z4hY/2,-HwX2H8Z4hY,2,text,3.9820311069488525,2.40418455898877,8,768,40,0.1,768,0.6164926290512085,ok,920a1052c09ebd1eedba6dd1ccfecd80f56f6f0fdcdd1f6dd32b0d90bcb89358
|
||||||
|
-HwX2H8Z4hY/2,-HwX2H8Z4hY,2,audio,3.9820311069488525,3.9192059241395896,392,74,40,0.1,74,0.9881895780563354,ok,920a1052c09ebd1eedba6dd1ccfecd80f56f6f0fdcdd1f6dd32b0d90bcb89358
|
||||||
|
-HwX2H8Z4hY/2,-HwX2H8Z4hY,2,vision,3.9820311069488525,0.8653643131256102,27,35,40,0.1,35,0.9814813733100891,ok,920a1052c09ebd1eedba6dd1ccfecd80f56f6f0fdcdd1f6dd32b0d90bcb89358
|
||||||
|
-HwX2H8Z4hY/5,-HwX2H8Z4hY,5,text,5.6529951095581055,3.0192474961280813,10,768,57,0.1,768,0.7021505832672119,ok,8d61b7f5840a8bd403d9bc6c3cd49029ffc4e304cf9c7834735c1a6bb5fd369d
|
||||||
|
-HwX2H8Z4hY/5,-HwX2H8Z4hY,5,audio,5.6529951095581055,5.595076320642555,555,74,57,0.1,74,0.9900854229927063,ok,8d61b7f5840a8bd403d9bc6c3cd49029ffc4e304cf9c7834735c1a6bb5fd369d
|
||||||
|
-HwX2H8Z4hY/5,-HwX2H8Z4hY,5,vision,5.6529951095581055,1.199999868869782,44,35,57,0.1,35,0.666666567325592,ok,8d61b7f5840a8bd403d9bc6c3cd49029ffc4e304cf9c7834735c1a6bb5fd369d
|
||||||
|
-HwX2H8Z4hY/6,-HwX2H8Z4hY,6,text,3.3580079078674316,1.3907382175326355,7,768,34,0.1,768,0.5562952756881714,ok,696576dcbd0dc42cb678fa40d8a0f0f419a3072bfa7662d653756ab29bba4c46
|
||||||
|
-HwX2H8Z4hY/6,-HwX2H8Z4hY,6,audio,3.3580079078674316,3.306980685486987,326,74,34,0.1,74,0.9901678562164307,ok,696576dcbd0dc42cb678fa40d8a0f0f419a3072bfa7662d653756ab29bba4c46
|
||||||
|
-HwX2H8Z4hY/6,-HwX2H8Z4hY,6,vision,3.3580079078674316,1.6413411736488344,21,35,34,0.1,35,0.841666579246521,ok,696576dcbd0dc42cb678fa40d8a0f0f419a3072bfa7662d653756ab29bba4c46
|
||||||
|
-HwX2H8Z4hY/9,-HwX2H8Z4hY,9,text,7.733983993530273,4.492685443162919,22,768,78,0.1,768,0.6705501675605774,ok,90c18e627f27a12c14e21ba24c79edf3008762b4ac26874b490ad3f888381702
|
||||||
|
-HwX2H8Z4hY/9,-HwX2H8Z4hY,9,audio,7.733983993530273,7.691443904670509,764,74,78,0.1,74,0.9946621060371399,ok,90c18e627f27a12c14e21ba24c79edf3008762b4ac26874b490ad3f888381702
|
||||||
|
-HwX2H8Z4hY/9,-HwX2H8Z4hY,9,vision,7.733983993530273,0.0,65,35,78,0.1,35,0.0,ok,90c18e627f27a12c14e21ba24c79edf3008762b4ac26874b490ad3f888381702
|
||||||
|
-NFrJFQijFE/1,-NFrJFQijFE,1,text,5.745999813079834,4.074365028738975,16,768,58,0.1,768,0.7835317254066467,ok,d8fcf16c6eb51c56947ec3d0549c6f3bc7a0cf1fc5f5d89e8b38d4249a08db4f
|
||||||
|
-NFrJFQijFE/1,-NFrJFQijFE,1,audio,5.745999813079834,5.699824074635634,566,74,58,0.1,74,0.9922494292259216,ok,d8fcf16c6eb51c56947ec3d0549c6f3bc7a0cf1fc5f5d89e8b38d4249a08db4f
|
||||||
|
-NFrJFQijFE/1,-NFrJFQijFE,1,vision,5.745999813079834,0.0,44,35,58,0.1,35,0.0,ok,d8fcf16c6eb51c56947ec3d0549c6f3bc7a0cf1fc5f5d89e8b38d4249a08db4f
|
||||||
|
-NFrJFQijFE/2,-NFrJFQijFE,2,text,6.855999946594238,3.4186642885208136,18,768,69,0.1,768,0.7431879639625549,ok,a54a5c144a64f5abd26b19aafa6cea8142ee08e9385fa508fa70c934f6b616c1
|
||||||
|
-NFrJFQijFE/2,-NFrJFQijFE,2,audio,6.855999946594238,6.798472889210726,670,74,69,0.1,74,0.9926568269729614,ok,a54a5c144a64f5abd26b19aafa6cea8142ee08e9385fa508fa70c934f6b616c1
|
||||||
|
-NFrJFQijFE/2,-NFrJFQijFE,2,vision,6.855999946594238,0.0,55,35,69,0.1,35,0.0,ok,a54a5c144a64f5abd26b19aafa6cea8142ee08e9385fa508fa70c934f6b616c1
|
||||||
|
-THoVjtIkeU/12,-THoVjtIkeU,12,text,14.896029472351074,7.941008102986966,39,768,149,0.1,768,0.6352806687355042,ok,e2f148f4f2e74724dc13b3ce870a0ce6a06740cfc224f5a9b3d0a870b196e7d4
|
||||||
|
-THoVjtIkeU/12,-THoVjtIkeU,12,audio,14.896029472351074,14.759312417861578,1481,74,149,0.1,74,0.9911767244338989,ok,e2f148f4f2e74724dc13b3ce870a0ce6a06740cfc224f5a9b3d0a870b196e7d4
|
||||||
|
-THoVjtIkeU/12,-THoVjtIkeU,12,vision,14.896029472351074,14.896029472351074,106,35,149,0.1,35,1.0,ok,e2f148f4f2e74724dc13b3ce870a0ce6a06740cfc224f5a9b3d0a870b196e7d4
|
||||||
|
-THoVjtIkeU/2,-THoVjtIkeU,2,text,4.2919921875,2.5793988468125453,10,768,43,0.1,768,0.7164996266365051,ok,35a6c2464ffd68dd96411edb5dc2763f0efb59f264c1d8b737e616fa6b0d1d5b
|
||||||
|
-THoVjtIkeU/2,-THoVjtIkeU,2,audio,4.2919921875,4.247343715932063,424,74,43,0.1,74,0.9895758628845215,ok,35a6c2464ffd68dd96411edb5dc2763f0efb59f264c1d8b737e616fa6b0d1d5b
|
||||||
|
-THoVjtIkeU/2,-THoVjtIkeU,2,vision,4.2919921875,4.2919921875,31,35,43,0.1,35,1.0,ok,35a6c2464ffd68dd96411edb5dc2763f0efb59f264c1d8b737e616fa6b0d1d5b
|
||||||
|
-THoVjtIkeU/6,-THoVjtIkeU,6,text,8.097004890441895,5.15853084437549,30,768,81,0.1,768,0.6787540316581726,ok,efaa55ee6394032c867046690a92c29edfd825005f0a090cca5ccb763b227f64
|
||||||
|
-THoVjtIkeU/6,-THoVjtIkeU,6,audio,8.097004890441895,8.027356093840018,799,74,81,0.1,74,0.9927164316177368,ok,efaa55ee6394032c867046690a92c29edfd825005f0a090cca5ccb763b227f64
|
||||||
|
-THoVjtIkeU/6,-THoVjtIkeU,6,vision,8.097004890441895,8.097004890441895,68,35,81,0.1,35,1.0,ok,efaa55ee6394032c867046690a92c29edfd825005f0a090cca5ccb763b227f64
|
||||||
|
-UuX1xuaiiE/1,-UuX1xuaiiE,1,text,10.350000381469727,5.7645069250837,27,768,104,0.1,768,0.6550575494766235,ok,d5bbda38fcec15817d2b87bab5dcc559d6d425f7d28a82e6c32d13f14b48650c
|
||||||
|
-UuX1xuaiiE/1,-UuX1xuaiiE,1,audio,10.350000381469727,10.268040973753543,1027,74,104,0.1,74,0.9925051927566528,ok,d5bbda38fcec15817d2b87bab5dcc559d6d425f7d28a82e6c32d13f14b48650c
|
||||||
|
-UuX1xuaiiE/1,-UuX1xuaiiE,1,vision,10.350000381469727,0.45000076293945285,83,35,104,0.1,35,0.8333339691162109,ok,d5bbda38fcec15817d2b87bab5dcc559d6d425f7d28a82e6c32d13f14b48650c
|
||||||
|
-UuX1xuaiiE/0,-UuX1xuaiiE,0,text,3.6333329677581787,2.204757870733737,9,768,37,0.1,768,0.6484582424163818,ok,f578b305d53bc152702f5e771de0fcd12d1e9aaefc5cefce3d2082a417aab7c8
|
||||||
|
-UuX1xuaiiE/0,-UuX1xuaiiE,0,audio,3.6333329677581787,3.6124320519937045,358,74,37,0.1,74,0.994590699672699,ok,f578b305d53bc152702f5e771de0fcd12d1e9aaefc5cefce3d2082a417aab7c8
|
||||||
|
-UuX1xuaiiE/0,-UuX1xuaiiE,0,vision,3.6333329677581787,1.700000047683716,25,35,37,0.1,35,0.9444444179534912,ok,f578b305d53bc152702f5e771de0fcd12d1e9aaefc5cefce3d2082a417aab7c8
|
||||||
|
-UuX1xuaiiE/3,-UuX1xuaiiE,3,text,4.1300129890441895,2.0997684210538865,9,768,42,0.1,768,0.7240580320358276,ok,eddb408f25bdc13e6de2fb74caeed709cb7a8f00640e7897135330437f66a2aa
|
||||||
|
-UuX1xuaiiE/3,-UuX1xuaiiE,3,audio,4.1300129890441895,4.093458598368876,404,74,42,0.1,74,0.992123007774353,ok,eddb408f25bdc13e6de2fb74caeed709cb7a8f00640e7897135330437f66a2aa
|
||||||
|
-UuX1xuaiiE/3,-UuX1xuaiiE,3,vision,4.1300129890441895,4.03001308441162,29,35,42,0.1,35,0.976190447807312,ok,eddb408f25bdc13e6de2fb74caeed709cb7a8f00640e7897135330437f66a2aa
|
||||||
|
-UuX1xuaiiE/6,-UuX1xuaiiE,6,text,8.113997459411621,4.803536714613438,22,768,82,0.1,768,0.7064024806022644,ok,f0131ac00e44410fbd32c547a0d421c2791172434fc0203bb969abe14a530532
|
||||||
|
-UuX1xuaiiE/6,-UuX1xuaiiE,6,audio,8.113997459411621,7.968213647201255,795,74,82,0.1,74,0.984533965587616,ok,f0131ac00e44410fbd32c547a0d421c2791172434fc0203bb969abe14a530532
|
||||||
|
-UuX1xuaiiE/6,-UuX1xuaiiE,6,vision,8.113997459411621,7.9306641817092896,68,35,82,0.1,35,0.9776423573493958,ok,f0131ac00e44410fbd32c547a0d421c2791172434fc0203bb969abe14a530532
|
||||||
|
-a55Q6RWvTA/3,-a55Q6RWvTA,3,text,22.15397071838379,13.423765965458015,65,768,222,0.1,768,0.6580277681350708,ok,15d029fc15f50b268b98f1e8abc65e45582e638c13a018d7aa74a18373275946
|
||||||
|
-a55Q6RWvTA/3,-a55Q6RWvTA,3,audio,22.15397071838379,21.95386351814141,2209,74,222,0.1,74,0.9912921786308289,ok,15d029fc15f50b268b98f1e8abc65e45582e638c13a018d7aa74a18373275946
|
||||||
|
-a55Q6RWvTA/3,-a55Q6RWvTA,3,vision,22.15397071838379,22.15397071838379,142,35,222,0.1,35,1.0,ok,15d029fc15f50b268b98f1e8abc65e45582e638c13a018d7aa74a18373275946
|
||||||
|
-aNfi7CP8vM/7,-aNfi7CP8vM,7,text,8.694987297058105,5.3913327027112254,18,768,87,0.1,768,0.7001731395721436,ok,0b6389f45bb966113c65a11998c6e7facdd24c8c42acf2e8cf32e0f2139402e1
|
||||||
|
-aNfi7CP8vM/7,-aNfi7CP8vM,7,audio,8.694987297058105,8.639379485394505,854,74,87,0.1,74,0.9945195913314819,ok,0b6389f45bb966113c65a11998c6e7facdd24c8c42acf2e8cf32e0f2139402e1
|
||||||
|
-aNfi7CP8vM/7,-aNfi7CP8vM,7,vision,8.694987297058105,8.694987297058105,74,35,87,0.1,35,1.0,ok,0b6389f45bb966113c65a11998c6e7facdd24c8c42acf2e8cf32e0f2139402e1
|
||||||
|
-aqamKhZ1Ec/0,-aqamKhZ1Ec,0,text,10.966667175292969,5.40069461874664,14,768,110,0.1,768,0.71061772108078,ok,6f6e674bbba4353399e9e7825217f37e6d8ca1adf9675f33cf37106f31c55548
|
||||||
|
-aqamKhZ1Ec/0,-aqamKhZ1Ec,0,audio,10.966667175292969,10.855676183829436,1090,74,110,0.1,74,0.9910455346107483,ok,6f6e674bbba4353399e9e7825217f37e6d8ca1adf9675f33cf37106f31c55548
|
||||||
|
-aqamKhZ1Ec/0,-aqamKhZ1Ec,0,vision,10.966667175292969,10.966667175292969,86,35,110,0.1,35,1.0,ok,6f6e674bbba4353399e9e7825217f37e6d8ca1adf9675f33cf37106f31c55548
|
||||||
|
-dxfTGcXJoc/1,-dxfTGcXJoc,1,text,16.16100311279297,9.47002421617508,36,768,162,0.1,768,0.6812966465950012,ok,b5ffdc98a4a98b8f0aae55dee5fed66a43bcb129f47a250a004c1cc74e572595
|
||||||
|
-dxfTGcXJoc/1,-dxfTGcXJoc,1,audio,16.16100311279297,15.984151583668348,1602,74,162,0.1,74,0.9900091290473938,ok,b5ffdc98a4a98b8f0aae55dee5fed66a43bcb129f47a250a004c1cc74e572595
|
||||||
|
-dxfTGcXJoc/1,-dxfTGcXJoc,1,vision,16.16100311279297,16.16100311279297,112,35,162,0.1,35,1.0,ok,b5ffdc98a4a98b8f0aae55dee5fed66a43bcb129f47a250a004c1cc74e572595
|
||||||
|
-dxfTGcXJoc/0,-dxfTGcXJoc,0,text,20.5,13.566675072815258,41,768,205,0.1,768,0.7333337664604187,ok,04bd0be907fb46c613f926fe1b06bc2c3b30881c293156510baa34a4105ade55
|
||||||
|
-dxfTGcXJoc/0,-dxfTGcXJoc,0,audio,20.5,20.265135636200775,2042,74,205,0.1,74,0.9891952872276306,ok,04bd0be907fb46c613f926fe1b06bc2c3b30881c293156510baa34a4105ade55
|
||||||
|
-dxfTGcXJoc/0,-dxfTGcXJoc,0,vision,20.5,15.983333587646484,134,35,205,0.1,35,0.99895840883255,ok,04bd0be907fb46c613f926fe1b06bc2c3b30881c293156510baa34a4105ade55
|
||||||
|
-dxfTGcXJoc/2,-dxfTGcXJoc,2,text,16.697982788085938,9.495353608578453,34,768,167,0.1,768,0.6593995094299316,ok,46a5e523b5dd00364c99a3bcbc6fae4c80e1382a84b5b5e681c46eb074ed10f7
|
||||||
|
-dxfTGcXJoc/2,-dxfTGcXJoc,2,audio,16.697982788085938,16.517766948940555,1664,74,167,0.1,74,0.9901677966117859,ok,46a5e523b5dd00364c99a3bcbc6fae4c80e1382a84b5b5e681c46eb074ed10f7
|
||||||
|
-dxfTGcXJoc/2,-dxfTGcXJoc,2,vision,16.697982788085938,16.697982788085938,115,35,167,0.1,35,1.0,ok,46a5e523b5dd00364c99a3bcbc6fae4c80e1382a84b5b5e681c46eb074ed10f7
|
||||||
|
-dxfTGcXJoc/6,-dxfTGcXJoc,6,text,12.735026359558105,8.155348146427425,27,768,128,0.1,768,0.715381383895874,ok,55c1ff86729ab21da743a5107ed07b220e41b4193cd1c0ba9f009e65c3ab0513
|
||||||
|
-dxfTGcXJoc/6,-dxfTGcXJoc,6,audio,12.735026359558105,12.576174248392515,1263,74,128,0.1,74,0.9889141917228699,ok,55c1ff86729ab21da743a5107ed07b220e41b4193cd1c0ba9f009e65c3ab0513
|
||||||
|
-dxfTGcXJoc/6,-dxfTGcXJoc,6,vision,12.735026359558105,12.5,95,35,128,0.1,35,1.0,ok,55c1ff86729ab21da743a5107ed07b220e41b4193cd1c0ba9f009e65c3ab0513
|
||||||
|
-egA8-b7-3M/26,-egA8-b7-3M,26,text,6.030990123748779,3.6452820789068947,15,768,61,0.1,768,0.6877890825271606,ok,11f78aed680ea9c5862a15e04ff6c1f8776c46d29c98b68d28773ff39d959685
|
||||||
|
-egA8-b7-3M/26,-egA8-b7-3M,26,audio,6.030990123748779,5.994409297769134,593,74,61,0.1,74,0.9945786595344543,ok,11f78aed680ea9c5862a15e04ff6c1f8776c46d29c98b68d28773ff39d959685
|
||||||
|
-egA8-b7-3M/26,-egA8-b7-3M,26,vision,6.030990123748779,6.030990123748779,47,35,61,0.1,35,1.0,ok,11f78aed680ea9c5862a15e04ff6c1f8776c46d29c98b68d28773ff39d959685
|
||||||
|
-egA8-b7-3M/17,-egA8-b7-3M,17,text,7.271028995513916,3.4275580286979666,16,768,73,0.1,768,0.7292676568031311,ok,f967e81e0edd51c3700671721afb3fbecbbc32652f27d2db40ea444e8d80c68f
|
||||||
|
-egA8-b7-3M/17,-egA8-b7-3M,17,audio,7.271028995513916,7.202622791158187,710,74,73,0.1,74,0.9916234016418457,ok,f967e81e0edd51c3700671721afb3fbecbbc32652f27d2db40ea444e8d80c68f
|
||||||
|
-egA8-b7-3M/17,-egA8-b7-3M,17,vision,7.271028995513916,7.271028995513916,60,35,73,0.1,35,1.0,ok,f967e81e0edd51c3700671721afb3fbecbbc32652f27d2db40ea444e8d80c68f
|
||||||
|
-egA8-b7-3M/18,-egA8-b7-3M,18,text,8.386002540588379,3.9958887480199343,22,768,84,0.1,768,0.6243576407432556,ok,fd4bc95a6adfb16a9588b7ed65cc2812186ce3d607804dc76f4486312951ef90
|
||||||
|
-egA8-b7-3M/18,-egA8-b7-3M,18,audio,8.386002540588379,8.326583495091747,827,74,84,0.1,74,0.9932008385658264,ok,fd4bc95a6adfb16a9588b7ed65cc2812186ce3d607804dc76f4486312951ef90
|
||||||
|
-egA8-b7-3M/18,-egA8-b7-3M,18,vision,8.386002540588379,8.386002540588379,71,35,84,0.1,35,1.0,ok,fd4bc95a6adfb16a9588b7ed65cc2812186ce3d607804dc76f4486312951ef90
|
||||||
|
-egA8-b7-3M/16,-egA8-b7-3M,16,text,5.538021087646484,3.3190146580338484,11,768,56,0.1,768,0.7375588417053223,ok,4c015f85e3bd901ecedde8c0c2ca4a756d9a77dfe4177625c481581f2037ff9c
|
||||||
|
-egA8-b7-3M/16,-egA8-b7-3M,16,audio,5.538021087646484,5.498074575251824,549,74,56,0.1,74,0.9940068125724792,ok,4c015f85e3bd901ecedde8c0c2ca4a756d9a77dfe4177625c481581f2037ff9c
|
||||||
|
-egA8-b7-3M/16,-egA8-b7-3M,16,vision,5.538021087646484,5.538021087646484,43,35,56,0.1,35,1.0,ok,4c015f85e3bd901ecedde8c0c2ca4a756d9a77dfe4177625c481581f2037ff9c
|
||||||
|
-egA8-b7-3M/13,-egA8-b7-3M,13,text,4.18398380279541,2.166685357689858,11,768,42,0.1,768,0.6372604370117188,ok,cbd627c225ebca37521ae238cb9ec254cad236822f4f37ebfbb039e6c56aea4c
|
||||||
|
-egA8-b7-3M/13,-egA8-b7-3M,13,audio,4.18398380279541,4.149822072886132,403,74,42,0.1,74,0.9924017786979675,ok,cbd627c225ebca37521ae238cb9ec254cad236822f4f37ebfbb039e6c56aea4c
|
||||||
|
-egA8-b7-3M/13,-egA8-b7-3M,13,vision,4.18398380279541,4.18398380279541,29,35,42,0.1,35,1.0,ok,cbd627c225ebca37521ae238cb9ec254cad236822f4f37ebfbb039e6c56aea4c
|
||||||
|
-egA8-b7-3M/1,-egA8-b7-3M,1,text,10.266016006469727,5.581577223725616,20,768,103,0.1,768,0.6490206122398376,ok,589a7989b1b4888568c31614a2d6beefba88fb87289c85b94df2d8051500403b
|
||||||
|
-egA8-b7-3M/1,-egA8-b7-3M,1,audio,10.266016006469727,10.199569694899225,1016,74,103,0.1,74,0.9937204718589783,ok,589a7989b1b4888568c31614a2d6beefba88fb87289c85b94df2d8051500403b
|
||||||
|
-egA8-b7-3M/1,-egA8-b7-3M,1,vision,10.266016006469727,10.266016006469727,83,35,103,0.1,35,1.0,ok,589a7989b1b4888568c31614a2d6beefba88fb87289c85b94df2d8051500403b
|
||||||
|
-egA8-b7-3M/6,-egA8-b7-3M,6,text,6.264974117279053,3.582069296762348,12,768,63,0.1,768,0.7164138555526733,ok,2e88a00d5e55863aa956b4b0949fc0e4f2975ccfd83b3194957f6fca86ecfce0
|
||||||
|
-egA8-b7-3M/6,-egA8-b7-3M,6,audio,6.264974117279053,6.2242991206613745,619,74,63,0.1,74,0.9942464828491211,ok,2e88a00d5e55863aa956b4b0949fc0e4f2975ccfd83b3194957f6fca86ecfce0
|
||||||
|
-egA8-b7-3M/6,-egA8-b7-3M,6,vision,6.264974117279053,6.264974117279053,50,35,63,0.1,35,1.0,ok,2e88a00d5e55863aa956b4b0949fc0e4f2975ccfd83b3194957f6fca86ecfce0
|
||||||
|
-egA8-b7-3M/9,-egA8-b7-3M,9,text,5.0899739265441895,2.328742109239101,10,768,51,0.1,768,0.597113311290741,ok,bcd643f328616312eae9b28e6c9aabbe8acaec3e4c4be58daeb45e3eee4d973c
|
||||||
|
-egA8-b7-3M/9,-egA8-b7-3M,9,audio,5.0899739265441895,5.05666382530251,494,74,51,0.1,74,0.9939422011375427,ok,bcd643f328616312eae9b28e6c9aabbe8acaec3e4c4be58daeb45e3eee4d973c
|
||||||
|
-egA8-b7-3M/9,-egA8-b7-3M,9,vision,5.0899739265441895,5.0899739265441895,38,35,51,0.1,35,1.0,ok,bcd643f328616312eae9b28e6c9aabbe8acaec3e4c4be58daeb45e3eee4d973c
|
||||||
|
-egA8-b7-3M/20,-egA8-b7-3M,20,text,6.644987106323242,3.9119302723556753,20,768,67,0.1,768,0.6985589861869812,ok,719ef133920e74b3ea33925d2bca0270b062199116867568f88ac83cadcd74e8
|
||||||
|
-egA8-b7-3M/20,-egA8-b7-3M,20,audio,6.644987106323242,6.588095611333847,648,74,67,0.1,74,0.992087185382843,ok,719ef133920e74b3ea33925d2bca0270b062199116867568f88ac83cadcd74e8
|
||||||
|
-egA8-b7-3M/20,-egA8-b7-3M,20,vision,6.644987106323242,6.644987106323242,54,35,67,0.1,35,1.0,ok,719ef133920e74b3ea33925d2bca0270b062199116867568f88ac83cadcd74e8
|
||||||
|
-iRBcNs9oI8/3,-iRBcNs9oI8,3,text,5.620999813079834,2.037480006366969,8,768,57,0.1,768,0.6791599988937378,ok,39c547acd1a8da2ebc6c6ab008191a5676420398a974cd9611cad087f0ccec50
|
||||||
|
-iRBcNs9oI8/3,-iRBcNs9oI8,3,audio,5.620999813079834,5.578905152469068,554,74,57,0.1,74,0.9954468607902527,ok,39c547acd1a8da2ebc6c6ab008191a5676420398a974cd9611cad087f0ccec50
|
||||||
|
-iRBcNs9oI8/3,-iRBcNs9oI8,3,vision,5.620999813079834,0.0,28,35,57,0.1,35,0.0,ok,39c547acd1a8da2ebc6c6ab008191a5676420398a974cd9611cad087f0ccec50
|
||||||
|
-iRBcNs9oI8/7,-iRBcNs9oI8,7,text,4.104000091552734,2.2118893228471284,9,768,42,0.1,768,0.7135127186775208,ok,cc7b1c7a06ec41d72007b5c42fb8036f30b66d5cf76e3f093160c22fc5e38081
|
||||||
|
-iRBcNs9oI8/7,-iRBcNs9oI8,7,audio,4.104000091552734,4.083554145129952,401,74,42,0.1,74,0.995795726776123,ok,cc7b1c7a06ec41d72007b5c42fb8036f30b66d5cf76e3f093160c22fc5e38081
|
||||||
|
-iRBcNs9oI8/7,-iRBcNs9oI8,7,vision,4.104000091552734,0.0,21,35,42,0.1,35,0.0,ok,cc7b1c7a06ec41d72007b5c42fb8036f30b66d5cf76e3f093160c22fc5e38081
|
||||||
|
-iRBcNs9oI8/6,-iRBcNs9oI8,6,text,2.9030001163482666,1.5899996414780622,5,768,30,0.1,768,0.7227271199226379,ok,030b0b8d8d4ae47799432ab71b66fcc881f077c5b0feef892c1f146dc0a592e3
|
||||||
|
-iRBcNs9oI8/6,-iRBcNs9oI8,6,audio,2.9030001163482666,2.892310954429008,283,74,30,0.1,74,0.9973600506782532,ok,030b0b8d8d4ae47799432ab71b66fcc881f077c5b0feef892c1f146dc0a592e3
|
||||||
|
-iRBcNs9oI8/6,-iRBcNs9oI8,6,vision,2.9030001163482666,0.0,15,35,30,0.1,35,0.0,ok,030b0b8d8d4ae47799432ab71b66fcc881f077c5b0feef892c1f146dc0a592e3
|
||||||
|
-iRBcNs9oI8/9,-iRBcNs9oI8,9,text,3.4159998893737793,1.094375005364418,5,768,35,0.1,768,0.643750011920929,ok,352bdcc73d3622d6abac1b09f2039b211d15b9bf7acee66f0b07783c9b7a4c74
|
||||||
|
-iRBcNs9oI8/9,-iRBcNs9oI8,9,audio,3.4159998893737793,3.383608031917263,338,74,35,0.1,74,0.993934154510498,ok,352bdcc73d3622d6abac1b09f2039b211d15b9bf7acee66f0b07783c9b7a4c74
|
||||||
|
-iRBcNs9oI8/9,-iRBcNs9oI8,9,vision,3.4159998893737793,0.0,17,35,35,0.1,35,0.0,ok,352bdcc73d3622d6abac1b09f2039b211d15b9bf7acee66f0b07783c9b7a4c74
|
||||||
|
-iRBcNs9oI8/8,-iRBcNs9oI8,8,text,8.093000411987305,3.901277539858493,21,768,81,0.1,768,0.6192578673362732,ok,72a49e3da2862addcde033bd2fbae957d086533664ca4dcbccd3a851bcd50412
|
||||||
|
-iRBcNs9oI8/8,-iRBcNs9oI8,8,audio,8.093000411987305,8.037513972489212,803,74,81,0.1,74,0.9947929382324219,ok,72a49e3da2862addcde033bd2fbae957d086533664ca4dcbccd3a851bcd50412
|
||||||
|
-iRBcNs9oI8/8,-iRBcNs9oI8,8,vision,8.093000411987305,0.0,41,35,81,0.1,35,0.0,ok,72a49e3da2862addcde033bd2fbae957d086533664ca4dcbccd3a851bcd50412
|
||||||
|
-lzEya4AM_4/5,-lzEya4AM_4,5,text,7.675000190734863,4.456424816604702,19,768,77,0.1,768,0.7427374720573425,ok,3b74de8d0e66fd754594f05985e5af3480adcc03555f79225d3b727cb4bed043
|
||||||
|
-lzEya4AM_4/5,-lzEya4AM_4,5,audio,7.675000190734863,7.5937163826178855,757,74,77,0.1,74,0.9901386499404907,ok,3b74de8d0e66fd754594f05985e5af3480adcc03555f79225d3b727cb4bed043
|
||||||
|
-lzEya4AM_4/5,-lzEya4AM_4,5,vision,7.675000190734863,7.675000190734863,64,35,77,0.1,35,1.0,ok,3b74de8d0e66fd754594f05985e5af3480adcc03555f79225d3b727cb4bed043
|
||||||
|
-lzEya4AM_4/6,-lzEya4AM_4,6,text,14.7919921875,9.074837291240696,43,768,148,0.1,768,0.7318416833877563,ok,ff5da5afb551001d58100ed2471caea51c741046d5396957efbf40082307ca18
|
||||||
|
-lzEya4AM_4/6,-lzEya4AM_4,6,audio,14.7919921875,14.634100490929308,1475,74,148,0.1,74,0.9907644391059875,ok,ff5da5afb551001d58100ed2471caea51c741046d5396957efbf40082307ca18
|
||||||
|
-lzEya4AM_4/6,-lzEya4AM_4,6,vision,14.7919921875,14.7919921875,105,35,148,0.1,35,1.0,ok,ff5da5afb551001d58100ed2471caea51c741046d5396957efbf40082307ca18
|
||||||
|
-mJ2ud6oKI8/1,-mJ2ud6oKI8,1,text,6.103000164031982,1.9617638647556308,13,768,62,0.1,768,0.8174015879631042,ok,c8e3acb4ca08679796c4c8cdfcd8c2b3cffc0465015efaffa1eccdb55ff4f45b
|
||||||
|
-mJ2ud6oKI8/1,-mJ2ud6oKI8,1,audio,6.103000164031982,5.938054213652739,606,74,62,0.1,74,1.0,ok,c8e3acb4ca08679796c4c8cdfcd8c2b3cffc0465015efaffa1eccdb55ff4f45b
|
||||||
|
-mJ2ud6oKI8/1,-mJ2ud6oKI8,1,vision,6.103000164031982,0.0,31,35,62,0.1,35,0.0,ok,c8e3acb4ca08679796c4c8cdfcd8c2b3cffc0465015efaffa1eccdb55ff4f45b
|
||||||
|
-mJ2ud6oKI8/2,-mJ2ud6oKI8,2,text,5.730999946594238,3.5213801980018595,11,768,58,0.1,768,0.9266790151596069,ok,c41790f8b75f886bd1f5d6d72cbab1985306d0b715e5ae47203f3d0c8a3e7b28
|
||||||
|
-mJ2ud6oKI8/2,-mJ2ud6oKI8,2,audio,5.730999946594238,5.5761080561457455,571,74,58,0.1,74,1.0,ok,c41790f8b75f886bd1f5d6d72cbab1985306d0b715e5ae47203f3d0c8a3e7b28
|
||||||
|
-mJ2ud6oKI8/2,-mJ2ud6oKI8,2,vision,5.730999946594238,0.0,29,35,58,0.1,35,0.0,ok,c41790f8b75f886bd1f5d6d72cbab1985306d0b715e5ae47203f3d0c8a3e7b28
|
||||||
|
-mJ2ud6oKI8/6,-mJ2ud6oKI8,6,text,2.256999969482422,1.1722640581429007,5,768,23,0.1,768,0.6512578129768372,ok,134a3a4ebbc423760133b0da0e90cfe84d75ea85df3968000afd84153c5d828d
|
||||||
|
-mJ2ud6oKI8/6,-mJ2ud6oKI8,6,audio,2.256999969482422,2.2208648198211596,223,74,23,0.1,74,0.9862568378448486,ok,134a3a4ebbc423760133b0da0e90cfe84d75ea85df3968000afd84153c5d828d
|
||||||
|
-mJ2ud6oKI8/6,-mJ2ud6oKI8,6,vision,2.256999969482422,2.256999969482422,12,35,23,0.1,35,1.0,ok,134a3a4ebbc423760133b0da0e90cfe84d75ea85df3968000afd84153c5d828d
|
||||||
|
-mJ2ud6oKI8/9,-mJ2ud6oKI8,9,text,7.0329999923706055,3.813062208145856,22,768,71,0.1,768,0.5777366757392883,ok,df3442f3ed8f894b507923e87e426ae9636790a4a502d00c350a215270cc3e47
|
||||||
|
-mJ2ud6oKI8/9,-mJ2ud6oKI8,9,audio,7.0329999923706055,6.9057026652870945,691,74,71,0.1,74,0.9846946001052856,ok,df3442f3ed8f894b507923e87e426ae9636790a4a502d00c350a215270cc3e47
|
||||||
|
-mJ2ud6oKI8/9,-mJ2ud6oKI8,9,vision,7.0329999923706055,6.832799911499023,35,35,71,0.1,35,0.9856857061386108,ok,df3442f3ed8f894b507923e87e426ae9636790a4a502d00c350a215270cc3e47
|
||||||
|
-mJ2ud6oKI8/8,-mJ2ud6oKI8,8,text,4.191999912261963,2.284389768779285,11,768,42,0.1,768,0.6529467701911926,ok,6af6614c9838417a48c4817f4dd5bf1e6669eb5f78f983e9a68adc113fcdc7cc
|
||||||
|
-mJ2ud6oKI8/8,-mJ2ud6oKI8,8,audio,4.191999912261963,4.127891841835266,418,74,42,0.1,74,0.9846959114074707,ok,6af6614c9838417a48c4817f4dd5bf1e6669eb5f78f983e9a68adc113fcdc7cc
|
||||||
|
-mJ2ud6oKI8/8,-mJ2ud6oKI8,8,vision,4.191999912261963,4.191999912261963,21,35,42,0.1,35,1.0,ok,6af6614c9838417a48c4817f4dd5bf1e6669eb5f78f983e9a68adc113fcdc7cc
|
||||||
|
-mqbVkbCndg/0,-mqbVkbCndg,0,text,6.466667175292969,3.891069056466222,12,768,65,0.1,768,0.7074670791625977,ok,cdeba953bba30907814acf72f3e35582dec263eaa25bdd922dad7fafe02e4dda
|
||||||
|
-mqbVkbCndg/0,-mqbVkbCndg,0,audio,6.466667175292969,6.400540961445988,639,74,65,0.1,74,0.9905118346214294,ok,cdeba953bba30907814acf72f3e35582dec263eaa25bdd922dad7fafe02e4dda
|
||||||
|
-mqbVkbCndg/0,-mqbVkbCndg,0,vision,6.466667175292969,6.466667175292969,53,35,65,0.1,35,1.0,ok,cdeba953bba30907814acf72f3e35582dec263eaa25bdd922dad7fafe02e4dda
|
||||||
|
-t217m2on-s/2,-t217m2on-s,2,text,5.6860032081604,2.3235132679343233,12,768,57,0.1,768,0.6279765367507935,ok,6713a983112173204502dc5adbb886bc6c38554b4627698611931420c47663a8
|
||||||
|
-t217m2on-s/2,-t217m2on-s,2,audio,5.6860032081604,5.64178691047269,564,74,57,0.1,74,0.9959543347358704,ok,6713a983112173204502dc5adbb886bc6c38554b4627698611931420c47663a8
|
||||||
|
-t217m2on-s/2,-t217m2on-s,2,vision,5.6860032081604,5.6860032081604,45,35,57,0.1,35,1.0,ok,6713a983112173204502dc5adbb886bc6c38554b4627698611931420c47663a8
|
||||||
|
-t217m2on-s/7,-t217m2on-s,7,text,17.183008193969727,10.16458021551371,45,768,172,0.1,768,0.6867958903312683,ok,41c76b8a733ecb780d56348f55d00f0ba0c07f44f46e93044e62d876710180b7
|
||||||
|
-t217m2on-s/7,-t217m2on-s,7,audio,17.183008193969727,17.025629438258505,1703,74,172,0.1,74,0.9928514361381531,ok,41c76b8a733ecb780d56348f55d00f0ba0c07f44f46e93044e62d876710180b7
|
||||||
|
-t217m2on-s/7,-t217m2on-s,7,vision,17.183008193969727,17.183008193969727,117,35,172,0.1,35,1.0,ok,41c76b8a733ecb780d56348f55d00f0ba0c07f44f46e93044e62d876710180b7
|
||||||
|
-tANM6ETl_M/3,-tANM6ETl_M,3,text,6.758008003234863,3.281156235933303,22,768,68,0.1,768,0.5859207510948181,ok,6d3af064c060dad3816a9e1dfa00101faebd8b7da7d8ceea85b0bc0fca70abd6
|
||||||
|
-tANM6ETl_M/3,-tANM6ETl_M,3,audio,6.758008003234863,6.702588935075579,661,74,68,0.1,74,0.9936580061912537,ok,6d3af064c060dad3816a9e1dfa00101faebd8b7da7d8ceea85b0bc0fca70abd6
|
||||||
|
-tANM6ETl_M/3,-tANM6ETl_M,3,vision,6.758008003234863,6.5333333015441895,54,35,68,0.1,35,0.9898989200592041,ok,6d3af064c060dad3816a9e1dfa00101faebd8b7da7d8ceea85b0bc0fca70abd6
|
||||||
|
-tPCytz4rww/11,-tPCytz4rww,11,text,5.466991901397705,3.309338267147543,17,768,55,0.1,768,0.6753751635551453,ok,6dc09d02467baaef67594ff32cffb01f8ed1565cda4c5abd161b8261a51e2e5d
|
||||||
|
-tPCytz4rww/11,-tPCytz4rww,11,audio,5.466991901397705,5.402410984844775,538,74,55,0.1,74,0.9890678524971008,ok,6dc09d02467baaef67594ff32cffb01f8ed1565cda4c5abd161b8261a51e2e5d
|
||||||
|
-tPCytz4rww/11,-tPCytz4rww,11,vision,5.466991901397705,5.466991901397705,43,35,55,0.1,35,1.0,ok,6dc09d02467baaef67594ff32cffb01f8ed1565cda4c5abd161b8261a51e2e5d
|
||||||
|
-tPCytz4rww/10,-tPCytz4rww,10,text,4.788997173309326,2.2889054998755456,12,768,48,0.1,768,0.5449775457382202,ok,758183192cbd5f0f20823d26ece3ddee40dec0aadec35247b2495510a53f52f8
|
||||||
|
-tPCytz4rww/10,-tPCytz4rww,10,audio,4.788997173309326,4.71733509463233,468,74,48,0.1,74,0.9894654750823975,ok,758183192cbd5f0f20823d26ece3ddee40dec0aadec35247b2495510a53f52f8
|
||||||
|
-tPCytz4rww/10,-tPCytz4rww,10,vision,4.788997173309326,4.788997173309326,35,35,48,0.1,35,1.0,ok,758183192cbd5f0f20823d26ece3ddee40dec0aadec35247b2495510a53f52f8
|
||||||
|
-tPCytz4rww/12,-tPCytz4rww,12,text,11.863997459411621,6.21423171013594,26,768,119,0.1,768,0.654129683971405,ok,e3c7f9cd67ad2d20fef997696dc278361f97f027db58762e22104e5720d509d5
|
||||||
|
-tPCytz4rww/12,-tPCytz4rww,12,audio,11.863997459411621,11.690646183812941,1174,74,119,0.1,74,0.9875938892364502,ok,e3c7f9cd67ad2d20fef997696dc278361f97f027db58762e22104e5720d509d5
|
||||||
|
-tPCytz4rww/12,-tPCytz4rww,12,vision,11.863997459411621,11.863997459411621,91,35,119,0.1,35,1.0,ok,e3c7f9cd67ad2d20fef997696dc278361f97f027db58762e22104e5720d509d5
|
||||||
|
-tPCytz4rww/16,-tPCytz4rww,16,text,7.205989837646484,2.154985956847668,16,768,73,0.1,768,0.4489554166793823,ok,8b3f82f9628792eebaf8c227becc2ebeba59bb8abb668bcca86e07fb1a992a93
|
||||||
|
-tPCytz4rww/16,-tPCytz4rww,16,audio,7.205989837646484,7.1144090478484685,714,74,73,0.1,74,0.988937258720398,ok,8b3f82f9628792eebaf8c227becc2ebeba59bb8abb668bcca86e07fb1a992a93
|
||||||
|
-tPCytz4rww/16,-tPCytz4rww,16,vision,7.205989837646484,7.205989837646484,59,35,73,0.1,35,1.0,ok,8b3f82f9628792eebaf8c227becc2ebeba59bb8abb668bcca86e07fb1a992a93
|
||||||
|
-tPCytz4rww/18,-tPCytz4rww,18,text,6.644987106323242,3.4940090492367735,16,768,67,0.1,768,0.6719247698783875,ok,eece70451e9c37a4b5f375bf646f66e09113fe3467f35bf1720cc14d986c1776
|
||||||
|
-tPCytz4rww/18,-tPCytz4rww,18,audio,6.644987106323242,6.553973502565075,657,74,67,0.1,74,0.9872123599052429,ok,eece70451e9c37a4b5f375bf646f66e09113fe3467f35bf1720cc14d986c1776
|
||||||
|
-tPCytz4rww/18,-tPCytz4rww,18,vision,6.644987106323242,6.644987106323242,54,35,67,0.1,35,1.0,ok,eece70451e9c37a4b5f375bf646f66e09113fe3467f35bf1720cc14d986c1776
|
||||||
|
-vxjVxOeScU/4,-vxjVxOeScU,4,text,9.127017974853516,5.366417751088739,21,768,92,0.1,768,0.6465563774108887,ok,00282df9a314394f19d616d561bd1916d1d30dca4c814dfb43630ca164f4c719
|
||||||
|
-vxjVxOeScU/4,-vxjVxOeScU,4,audio,9.127017974853516,9.08216616914079,904,74,92,0.1,74,0.9960808753967285,ok,00282df9a314394f19d616d561bd1916d1d30dca4c814dfb43630ca164f4c719
|
||||||
|
-vxjVxOeScU/4,-vxjVxOeScU,4,vision,9.127017974853516,9.127017974853516,77,35,92,0.1,35,1.0,ok,00282df9a314394f19d616d561bd1916d1d30dca4c814dfb43630ca164f4c719
|
||||||
|
-wMB_hJL-3o/7,-wMB_hJL-3o,7,text,6.044010162353516,3.555890290439128,22,768,61,0.1,768,0.658498227596283,ok,04a73d73fd150b07edab8e652fd14c9cb4a0f6e017ea97b17e8146e377b9fae7
|
||||||
|
-wMB_hJL-3o/7,-wMB_hJL-3o,7,audio,6.044010162353516,5.968699067508853,598,74,61,0.1,74,0.9895980954170227,ok,04a73d73fd150b07edab8e652fd14c9cb4a0f6e017ea97b17e8146e377b9fae7
|
||||||
|
-wMB_hJL-3o/7,-wMB_hJL-3o,7,vision,6.044010162353516,5.944010257720948,48,35,61,0.1,35,0.9836065769195557,ok,04a73d73fd150b07edab8e652fd14c9cb4a0f6e017ea97b17e8146e377b9fae7
|
||||||
|
-wny0OAz3g8/1,-wny0OAz3g8,1,text,6.844009876251221,4.170902410242706,23,768,69,0.1,768,0.62252277135849,ok,c5535859f129ce04da5f5b5104d02d4168057b9d947a5349d423278ad4f3d8e8
|
||||||
|
-wny0OAz3g8/1,-wny0OAz3g8,1,audio,6.844009876251221,6.814914991726747,678,74,69,0.1,74,0.995954155921936,ok,c5535859f129ce04da5f5b5104d02d4168057b9d947a5349d423278ad4f3d8e8
|
||||||
|
-wny0OAz3g8/1,-wny0OAz3g8,1,vision,6.844009876251221,3.327343463897705,56,35,69,0.1,35,0.9950981140136719,ok,c5535859f129ce04da5f5b5104d02d4168057b9d947a5349d423278ad4f3d8e8
|
||||||
|
-wny0OAz3g8/0,-wny0OAz3g8,0,text,3.3333330154418945,1.6946923546493056,9,768,34,0.1,768,0.6052471995353699,ok,1cc21b4432e2f4b592284ee40558a18acc7d208019117aebdb2f6a1b67ea198d
|
||||||
|
-wny0OAz3g8/0,-wny0OAz3g8,0,audio,3.3333330154418945,3.3145942848276446,328,74,34,0.1,74,0.9955413937568665,ok,1cc21b4432e2f4b592284ee40558a18acc7d208019117aebdb2f6a1b67ea198d
|
||||||
|
-wny0OAz3g8/0,-wny0OAz3g8,0,vision,3.3333330154418945,2.6999998092651367,21,35,34,0.1,35,0.9999999403953552,ok,1cc21b4432e2f4b592284ee40558a18acc7d208019117aebdb2f6a1b67ea198d
|
||||||
|
-wny0OAz3g8/3,-wny0OAz3g8,3,text,6.146028995513916,4.208628869801758,20,768,62,0.1,768,0.7014381289482117,ok,ea917c506bc9fa39460f2b722f7c416f646312bb17cd74baa618905b54837bce
|
||||||
|
-wny0OAz3g8/3,-wny0OAz3g8,3,audio,6.146028995513916,6.13413711335208,610,74,62,0.1,74,0.9980819821357727,ok,ea917c506bc9fa39460f2b722f7c416f646312bb17cd74baa618905b54837bce
|
||||||
|
-wny0OAz3g8/3,-wny0OAz3g8,3,vision,6.146028995513916,1.5293622016906734,49,35,62,0.1,35,0.9895832538604736,ok,ea917c506bc9fa39460f2b722f7c416f646312bb17cd74baa618905b54837bce
|
||||||
|
-wny0OAz3g8/2,-wny0OAz3g8,2,text,6.686978816986084,4.909538650512694,23,768,67,0.1,768,0.7553136348724365,ok,1825fb37db2e914fb616cd80aebd613afd873eecb7edd4d0b53158bfda0e522a
|
||||||
|
-wny0OAz3g8/2,-wny0OAz3g8,2,audio,6.686978816986084,6.656249635122918,653,74,67,0.1,74,0.9957627058029175,ok,1825fb37db2e914fb616cd80aebd613afd873eecb7edd4d0b53158bfda0e522a
|
||||||
|
-wny0OAz3g8/2,-wny0OAz3g8,2,vision,6.686978816986084,5.4166669845581055,54,35,67,0.1,35,0.9848485589027405,ok,1825fb37db2e914fb616cd80aebd613afd873eecb7edd4d0b53158bfda0e522a
|
||||||
|
-wny0OAz3g8/5,-wny0OAz3g8,5,text,6.9119791984558105,3.941009076312184,20,768,70,0.1,768,0.6063091158866882,ok,6c6a8a4cba654b25190ab1793bfae81c3d2159ecd6cbc02d3fcbd598543acce0
|
||||||
|
-wny0OAz3g8/5,-wny0OAz3g8,5,audio,6.9119791984558105,6.879425679509704,675,74,70,0.1,74,0.9961636066436768,ok,6c6a8a4cba654b25190ab1793bfae81c3d2159ecd6cbc02d3fcbd598543acce0
|
||||||
|
-wny0OAz3g8/5,-wny0OAz3g8,5,vision,6.9119791984558105,1.4953122138977046,56,35,70,0.1,35,0.9895831346511841,ok,6c6a8a4cba654b25190ab1793bfae81c3d2159ecd6cbc02d3fcbd598543acce0
|
||||||
|
-wny0OAz3g8/7,-wny0OAz3g8,7,text,8.06796932220459,5.293590973317621,29,768,81,0.1,768,0.6700748205184937,ok,cbf9cac1048ce191aef781d0c4563eb97f9f3ddd52eb1f9551fa78d31d70e7cd
|
||||||
|
-wny0OAz3g8/7,-wny0OAz3g8,7,audio,8.06796932220459,8.030726903515895,792,74,81,0.1,74,0.9956274628639221,ok,cbf9cac1048ce191aef781d0c4563eb97f9f3ddd52eb1f9551fa78d31d70e7cd
|
||||||
|
-wny0OAz3g8/7,-wny0OAz3g8,7,vision,8.06796932220459,7.06796932220459,68,35,81,0.1,35,1.0,ok,cbf9cac1048ce191aef781d0c4563eb97f9f3ddd52eb1f9551fa78d31d70e7cd
|
||||||
|
-wny0OAz3g8/9,-wny0OAz3g8,9,text,6.2130208015441895,4.012809376046059,20,768,63,0.1,768,0.6688015460968018,ok,b76f70fe783ee76bf1b637137226c98e67d13a6b15ec725894175b36cd76a136
|
||||||
|
-wny0OAz3g8/9,-wny0OAz3g8,9,audio,6.2130208015441895,6.190979680177328,605,74,63,0.1,74,0.9969837069511414,ok,b76f70fe783ee76bf1b637137226c98e67d13a6b15ec725894175b36cd76a136
|
||||||
|
-wny0OAz3g8/9,-wny0OAz3g8,9,vision,6.2130208015441895,5.8130205273628235,49,35,63,0.1,35,0.9672130942344666,ok,b76f70fe783ee76bf1b637137226c98e67d13a6b15ec725894175b36cd76a136
|
||||||
|
-571d8cVauQ/0,-571d8cVauQ,0,text,5.066667079925537,3.21007038205862,12,768,51,0.1,768,0.6978414058685303,ok,defc62e1530d9b0cd27f2acfe042b71dd84d7e8c85f1df261f04912091976088
|
||||||
|
-571d8cVauQ/0,-571d8cVauQ,0,audio,5.066667079925537,5.013243622071034,500,74,51,0.1,74,0.9903978705406189,ok,defc62e1530d9b0cd27f2acfe042b71dd84d7e8c85f1df261f04912091976088
|
||||||
|
-571d8cVauQ/0,-571d8cVauQ,0,vision,5.066667079925537,4.9666670799255375,39,35,51,0.1,35,1.0,ok,defc62e1530d9b0cd27f2acfe042b71dd84d7e8c85f1df261f04912091976088
|
||||||
|
-571d8cVauQ/5,-571d8cVauQ,5,text,15.272981643676758,9.052198391593993,43,768,153,0.1,768,0.7017207741737366,ok,e304c03fac1f477823fe0926c4fb421b9c4e4091f85718b729afe856c24a7d3e
|
||||||
|
-571d8cVauQ/5,-571d8cVauQ,5,audio,15.272981643676758,15.142900936426344,1512,74,153,0.1,74,0.9918006062507629,ok,e304c03fac1f477823fe0926c4fb421b9c4e4091f85718b729afe856c24a7d3e
|
||||||
|
-571d8cVauQ/5,-571d8cVauQ,5,vision,15.272981643676758,15.272981643676758,108,35,153,0.1,35,1.0,ok,e304c03fac1f477823fe0926c4fb421b9c4e4091f85718b729afe856c24a7d3e
|
||||||
|
-I_e4mIh0yE/1,-I_e4mIh0yE,1,text,7.622004985809326,3.535097924340517,18,768,77,0.1,768,0.6546477675437927,ok,8453a966d4bd1f0179eb86a35cd34c0864a3a3bc2fa86f37783060cc0d14f5df
|
||||||
|
-I_e4mIh0yE/1,-I_e4mIh0yE,1,audio,7.622004985809326,7.497626475627357,745,74,77,0.1,74,0.985649049282074,ok,8453a966d4bd1f0179eb86a35cd34c0864a3a3bc2fa86f37783060cc0d14f5df
|
||||||
|
-I_e4mIh0yE/1,-I_e4mIh0yE,1,vision,7.622004985809326,7.622004985809326,63,35,77,0.1,35,1.0,ok,8453a966d4bd1f0179eb86a35cd34c0864a3a3bc2fa86f37783060cc0d14f5df
|
||||||
|
-I_e4mIh0yE/3,-I_e4mIh0yE,3,text,9.163021087646484,4.626299023628236,22,768,92,0.1,768,0.6251755356788635,ok,a015f02ae51d74c42ad27e2a0831263a303f0658ac8d984819e27ad61342b23e
|
||||||
|
-I_e4mIh0yE/3,-I_e4mIh0yE,3,audio,9.163021087646484,9.017804326318405,900,74,92,0.1,74,0.9852646589279175,ok,a015f02ae51d74c42ad27e2a0831263a303f0658ac8d984819e27ad61342b23e
|
||||||
|
-I_e4mIh0yE/3,-I_e4mIh0yE,3,vision,9.163021087646484,9.163021087646484,77,35,92,0.1,35,1.0,ok,a015f02ae51d74c42ad27e2a0831263a303f0658ac8d984819e27ad61342b23e
|
||||||
|
-UacrmKiTn4/10,-UacrmKiTn4,10,text,7.361979007720947,2.732618988305329,18,768,74,0.1,768,0.47940683364868164,ok,00312e0602dbeac970de5d7dd0e8a970e514fe45d7fe6625b4f454b299e36360
|
||||||
|
-UacrmKiTn4/10,-UacrmKiTn4,10,audio,7.361979007720947,7.319425493478775,725,74,74,0.1,74,0.9955651164054871,ok,00312e0602dbeac970de5d7dd0e8a970e514fe45d7fe6625b4f454b299e36360
|
||||||
|
-UacrmKiTn4/10,-UacrmKiTn4,10,vision,7.361979007720947,7.361979007720947,60,35,74,0.1,35,1.0,ok,00312e0602dbeac970de5d7dd0e8a970e514fe45d7fe6625b4f454b299e36360
|
||||||
|
-UacrmKiTn4/4,-UacrmKiTn4,4,text,4.741015911102295,3.0301611527800563,14,768,48,0.1,768,0.7214669585227966,ok,a7f2ab0d6ee2e2a2d5cd4239f03e6cac67c618ddb744891323df3ff87e29af84
|
||||||
|
-UacrmKiTn4/4,-UacrmKiTn4,4,audio,4.741015911102295,4.704772258610339,469,74,48,0.1,74,0.9926760792732239,ok,a7f2ab0d6ee2e2a2d5cd4239f03e6cac67c618ddb744891323df3ff87e29af84
|
||||||
|
-UacrmKiTn4/4,-UacrmKiTn4,4,vision,4.741015911102295,4.741015911102295,35,35,48,0.1,35,1.0,ok,a7f2ab0d6ee2e2a2d5cd4239f03e6cac67c618ddb744891323df3ff87e29af84
|
||||||
|
-hnBHBN8p5A/7,-hnBHBN8p5A,7,text,7.372000217437744,4.009185888338835,13,768,74,0.1,768,0.657243549823761,ok,0ec6fb76226315d3fafbb483e3500e60cc8a6729617b10e4b4b76163595080fb
|
||||||
|
-hnBHBN8p5A/7,-hnBHBN8p5A,7,audio,7.372000217437744,7.281608330236899,731,74,74,0.1,74,0.9905768632888794,ok,0ec6fb76226315d3fafbb483e3500e60cc8a6729617b10e4b4b76163595080fb
|
||||||
|
-hnBHBN8p5A/7,-hnBHBN8p5A,7,vision,7.372000217437744,0.0,37,35,74,0.1,35,0.0,ok,0ec6fb76226315d3fafbb483e3500e60cc8a6729617b10e4b4b76163595080fb
|
||||||
|
-hnBHBN8p5A/6,-hnBHBN8p5A,6,text,6.401000022888184,3.580871792882682,13,768,65,0.1,768,0.7618876099586487,ok,997f6808ac3973ddd71800e91c8e10755c10b5037186d84348a79e00d672fdc7
|
||||||
|
-hnBHBN8p5A/6,-hnBHBN8p5A,6,audio,6.401000022888184,6.346108104731585,634,74,65,0.1,74,0.9932082295417786,ok,997f6808ac3973ddd71800e91c8e10755c10b5037186d84348a79e00d672fdc7
|
||||||
|
-hnBHBN8p5A/6,-hnBHBN8p5A,6,vision,6.401000022888184,0.0,32,35,65,0.1,35,0.0,ok,997f6808ac3973ddd71800e91c8e10755c10b5037186d84348a79e00d672fdc7
|
||||||
|
-qDkUB0GgYY/6,-qDkUB0GgYY,6,text,4.677995204925537,2.927082145679743,16,768,47,0.1,768,0.6504627466201782,ok,ab19bf4a761d3919b7113b1104ccaadbbf0c4e7b45694904d094280a2c329c8b
|
||||||
|
-qDkUB0GgYY/6,-qDkUB0GgYY,6,audio,4.677995204925537,4.653116947730401,462,74,47,0.1,74,0.9946085214614868,ok,ab19bf4a761d3919b7113b1104ccaadbbf0c4e7b45694904d094280a2c329c8b
|
||||||
|
-qDkUB0GgYY/6,-qDkUB0GgYY,6,vision,4.677995204925537,4.677995204925537,35,35,47,0.1,35,1.0,ok,ab19bf4a761d3919b7113b1104ccaadbbf0c4e7b45694904d094280a2c329c8b
|
||||||
|
-uywlfIYOS8/4,-uywlfIYOS8,4,text,6.044987201690674,3.7997015411034236,18,768,61,0.1,768,0.6908547878265381,ok,fde46f5222cf5cfec7af9c93214c972ce2d686b0bf9c90a291021601968986f9
|
||||||
|
-uywlfIYOS8/4,-uywlfIYOS8,4,audio,6.044987201690674,6.015325370511492,589,74,61,0.1,74,0.9957761168479919,ok,fde46f5222cf5cfec7af9c93214c972ce2d686b0bf9c90a291021601968986f9
|
||||||
|
-uywlfIYOS8/4,-uywlfIYOS8,4,vision,6.044987201690674,6.044987201690674,48,35,61,0.1,35,1.0,ok,fde46f5222cf5cfec7af9c93214c972ce2d686b0bf9c90a291021601968986f9
|
||||||
|
-6rXp3zJ3kc/8,-6rXp3zJ3kc,8,text,12.805012702941895,6.537867438234387,34,768,129,0.1,768,0.63474440574646,ok,59b77a0f873b7fa95e7327b93d64bb5a4bdae84039321988ef70bba8e63ee1ae
|
||||||
|
-6rXp3zJ3kc/8,-6rXp3zJ3kc,8,audio,12.805012702941895,12.672918005408468,1274,74,129,0.1,74,0.9899758696556091,ok,59b77a0f873b7fa95e7327b93d64bb5a4bdae84039321988ef70bba8e63ee1ae
|
||||||
|
-6rXp3zJ3kc/8,-6rXp3zJ3kc,8,vision,12.805012702941895,12.805012702941895,95,35,129,0.1,35,1.0,ok,59b77a0f873b7fa95e7327b93d64bb5a4bdae84039321988ef70bba8e63ee1ae
|
||||||
|
-9y-fZ3swSY/0,-9y-fZ3swSY,0,text,6.8333330154418945,3.536603624001146,22,768,69,0.1,768,0.5440928339958191,ok,29a1fd181293cdbe24e50edcad51e3cc0f27a5824d840bcc39d453b272fb6e3c
|
||||||
|
-9y-fZ3swSY/0,-9y-fZ3swSY,0,audio,6.8333330154418945,6.794324037513217,676,74,69,0.1,74,0.9948647022247314,ok,29a1fd181293cdbe24e50edcad51e3cc0f27a5824d840bcc39d453b272fb6e3c
|
||||||
|
-9y-fZ3swSY/0,-9y-fZ3swSY,0,vision,6.8333330154418945,6.8333330154418945,57,35,69,0.1,35,1.0,ok,29a1fd181293cdbe24e50edcad51e3cc0f27a5824d840bcc39d453b272fb6e3c
|
||||||
|
-9y-fZ3swSY/4,-9y-fZ3swSY,4,text,2.8210289478302,1.160462707281113,9,768,29,0.1,768,0.5274830460548401,ok,fe04a26710e85c58bdaad5a48618c7e0f742d82b0d920eb4a046c8f14a6921dc
|
||||||
|
-9y-fZ3swSY/4,-9y-fZ3swSY,4,audio,2.8210289478302,2.7937038478819103,267,74,29,0.1,74,0.9926167726516724,ok,fe04a26710e85c58bdaad5a48618c7e0f742d82b0d920eb4a046c8f14a6921dc
|
||||||
|
-9y-fZ3swSY/4,-9y-fZ3swSY,4,vision,2.8210289478302,2.7210289478302006,15,35,29,0.1,35,1.0,ok,fe04a26710e85c58bdaad5a48618c7e0f742d82b0d920eb4a046c8f14a6921dc
|
||||||
|
-9y-fZ3swSY/8,-9y-fZ3swSY,8,text,4.988996982574463,2.1943973237220367,12,768,50,0.1,768,0.562736988067627,ok,6140d031e614b62a5d8df73346fef2417c28cee3501310aa95aa40d358b4b256
|
||||||
|
-9y-fZ3swSY/8,-9y-fZ3swSY,8,audio,4.988996982574463,4.9497673825665816,492,74,50,0.1,74,0.9931800365447998,ok,6140d031e614b62a5d8df73346fef2417c28cee3501310aa95aa40d358b4b256
|
||||||
|
-9y-fZ3swSY/8,-9y-fZ3swSY,8,vision,4.988996982574463,3.688997030258178,37,35,50,0.1,35,0.9487179517745972,ok,6140d031e614b62a5d8df73346fef2417c28cee3501310aa95aa40d358b4b256
|
||||||
|
-AUZQgSxyPQ/2,-AUZQgSxyPQ,2,text,22.511003494262695,12.134252551943066,41,768,226,0.1,768,0.6386448740959167,ok,e09034b30bb4f8fb9f9c2568afc20f2a764e327fe6ef4e13f21804eaac51bb7b
|
||||||
|
-AUZQgSxyPQ/2,-AUZQgSxyPQ,2,audio,22.511003494262695,22.31367928599184,2241,74,226,0.1,74,0.9917553663253784,ok,e09034b30bb4f8fb9f9c2568afc20f2a764e327fe6ef4e13f21804eaac51bb7b
|
||||||
|
-AUZQgSxyPQ/2,-AUZQgSxyPQ,2,vision,22.511003494262695,22.511003494262695,144,35,226,0.1,35,1.0,ok,e09034b30bb4f8fb9f9c2568afc20f2a764e327fe6ef4e13f21804eaac51bb7b
|
||||||
|
-HeZS2-Prhc/2,-HeZS2-Prhc,2,text,8.261979103088379,3.552112135011702,16,768,83,0.1,768,0.6830984950065613,ok,edb0fa22fdfbd990c8174764af78670bd2ecad799240d35d96a1124a3cd9dc0c
|
||||||
|
-HeZS2-Prhc/2,-HeZS2-Prhc,2,audio,8.261979103088379,8.205439110059995,816,74,83,0.1,74,0.9933875799179077,ok,edb0fa22fdfbd990c8174764af78670bd2ecad799240d35d96a1124a3cd9dc0c
|
||||||
|
-HeZS2-Prhc/2,-HeZS2-Prhc,2,vision,8.261979103088379,8.261979103088379,70,35,83,0.1,35,1.0,ok,edb0fa22fdfbd990c8174764af78670bd2ecad799240d35d96a1124a3cd9dc0c
|
||||||
|
-MeTTeMJBNc/0,-MeTTeMJBNc,0,text,9.300000190734863,4.416411150246859,21,768,94,0.1,768,0.6400595903396606,ok,54e2097f7bdb455d243e72254e5af7ed16160739f2f3cfd22f57d2729ec1acbe
|
||||||
|
-MeTTeMJBNc/0,-MeTTeMJBNc,0,audio,9.300000190734863,9.216486695006088,922,74,94,0.1,74,0.9939717650413513,ok,54e2097f7bdb455d243e72254e5af7ed16160739f2f3cfd22f57d2729ec1acbe
|
||||||
|
-MeTTeMJBNc/0,-MeTTeMJBNc,0,vision,9.300000190734863,9.200000190734862,78,35,94,0.1,35,1.0,ok,54e2097f7bdb455d243e72254e5af7ed16160739f2f3cfd22f57d2729ec1acbe
|
||||||
|
-MeTTeMJBNc/13,-MeTTeMJBNc,13,text,5.430013179779053,2.997564935684203,14,768,55,0.1,768,0.6661255359649658,ok,80c5d996e5d7b61c58e1fe32d706efcaf80c8c8cddb5c37a5c66d527786c2a44
|
||||||
|
-MeTTeMJBNc/13,-MeTTeMJBNc,13,audio,5.430013179779053,5.368323606413764,526,74,55,0.1,74,0.9918515682220459,ok,80c5d996e5d7b61c58e1fe32d706efcaf80c8c8cddb5c37a5c66d527786c2a44
|
||||||
|
-MeTTeMJBNc/13,-MeTTeMJBNc,13,vision,5.430013179779053,5.430013179779053,41,35,55,0.1,35,1.0,ok,80c5d996e5d7b61c58e1fe32d706efcaf80c8c8cddb5c37a5c66d527786c2a44
|
||||||
|
-MeTTeMJBNc/7,-MeTTeMJBNc,7,text,10.51699161529541,6.728291415888818,31,768,106,0.1,768,0.7313360571861267,ok,40716dc7402fd6398038eded58ce939ec76be6afbeea346b968bf4c427088f3a
|
||||||
|
-MeTTeMJBNc/7,-MeTTeMJBNc,7,audio,10.51699161529541,10.441329748243898,1042,74,106,0.1,74,0.9934103488922119,ok,40716dc7402fd6398038eded58ce939ec76be6afbeea346b968bf4c427088f3a
|
||||||
|
-MeTTeMJBNc/7,-MeTTeMJBNc,7,vision,10.51699161529541,10.51699161529541,84,35,106,0.1,35,1.0,ok,40716dc7402fd6398038eded58ce939ec76be6afbeea346b968bf4c427088f3a
|
||||||
|
-RfYyzHpjk4/11,-RfYyzHpjk4,11,text,5.238996982574463,3.062753976881504,17,768,53,0.1,768,0.6960803866386414,ok,96e481cd732001239639c8cb4e7f94e570925e5a15940920ffab0cec1f13f904
|
||||||
|
-RfYyzHpjk4/11,-RfYyzHpjk4,11,audio,5.238996982574463,5.202267333462432,508,74,53,0.1,74,0.9958056807518005,ok,96e481cd732001239639c8cb4e7f94e570925e5a15940920ffab0cec1f13f904
|
||||||
|
-RfYyzHpjk4/11,-RfYyzHpjk4,11,vision,5.238996982574463,5.238996982574463,40,35,53,0.1,35,1.0,ok,96e481cd732001239639c8cb4e7f94e570925e5a15940920ffab0cec1f13f904
|
||||||
|
-RfYyzHpjk4/8,-RfYyzHpjk4,8,text,4.588996887207031,2.8058770607668286,17,768,46,0.1,768,0.684955358505249,ok,ed8c672c532b8c7d1c830c82d6e2c14fcb3bbd3dc1a86670a03a6c8e83c02f3c
|
||||||
|
-RfYyzHpjk4/8,-RfYyzHpjk4,8,audio,4.588996887207031,4.553010454316745,454,74,46,0.1,74,0.994469404220581,ok,ed8c672c532b8c7d1c830c82d6e2c14fcb3bbd3dc1a86670a03a6c8e83c02f3c
|
||||||
|
-RfYyzHpjk4/8,-RfYyzHpjk4,8,vision,4.588996887207031,4.588996887207031,33,35,46,0.1,35,1.0,ok,ed8c672c532b8c7d1c830c82d6e2c14fcb3bbd3dc1a86670a03a6c8e83c02f3c
|
||||||
|
-RfYyzHpjk4/2,-RfYyzHpjk4,2,text,5.516016006469727,3.372718970850109,25,768,56,0.1,768,0.7175998091697693,ok,5e428ee22de58c9d6561b9b4353bef9dfb811a5841d9ab07b5bc453d697ad17f
|
||||||
|
-RfYyzHpjk4/2,-RfYyzHpjk4,2,audio,5.516016006469727,5.478826378849712,540,74,56,0.1,74,0.9948742389678955,ok,5e428ee22de58c9d6561b9b4353bef9dfb811a5841d9ab07b5bc453d697ad17f
|
||||||
|
-RfYyzHpjk4/2,-RfYyzHpjk4,2,vision,5.516016006469727,5.416016006469727,43,35,56,0.1,35,1.0,ok,5e428ee22de58c9d6561b9b4353bef9dfb811a5841d9ab07b5bc453d697ad17f
|
||||||
|
-UUCSKoHeMA/0,-UUCSKoHeMA,0,text,7.800000190734863,4.133793779369442,18,768,79,0.1,768,0.666740894317627,ok,3f9792888ec8e969f61cfb6cb4de7bdf7ef8944afe0a9d2a9d13586e1f13b898
|
||||||
|
-UUCSKoHeMA/0,-UUCSKoHeMA,0,audio,7.800000190734863,7.750270468963159,774,74,79,0.1,74,0.9943835139274597,ok,3f9792888ec8e969f61cfb6cb4de7bdf7ef8944afe0a9d2a9d13586e1f13b898
|
||||||
|
-UUCSKoHeMA/0,-UUCSKoHeMA,0,vision,7.800000190734863,7.700000190734862,66,35,79,0.1,35,1.0,ok,3f9792888ec8e969f61cfb6cb4de7bdf7ef8944afe0a9d2a9d13586e1f13b898
|
||||||
|
-ri04Z7vwnc/0,-ri04Z7vwnc,0,text,2.806999921798706,2.1131306469440463,8,768,29,0.1,768,0.7826409935951233,ok,93021c70c8fad20ad3fded80e7fc790a2f9c96835de58a3b694c6906de42684b
|
||||||
|
-ri04Z7vwnc/0,-ri04Z7vwnc,0,audio,2.806999921798706,2.8018783078000356,277,74,29,0.1,74,0.9982975125312805,ok,93021c70c8fad20ad3fded80e7fc790a2f9c96835de58a3b694c6906de42684b
|
||||||
|
-ri04Z7vwnc/0,-ri04Z7vwnc,0,vision,2.806999921798706,0.0,14,35,29,0.1,35,0.0,ok,93021c70c8fad20ad3fded80e7fc790a2f9c96835de58a3b694c6906de42684b
|
||||||
|
-ri04Z7vwnc/2,-ri04Z7vwnc,2,text,6.0269999504089355,3.9384737918153405,16,768,61,0.1,768,0.7431082725524902,ok,f5993ce3a6ee56298462c08692a96f78c7d914d7264257ed401de7d6e17fd148
|
||||||
|
-ri04Z7vwnc/2,-ri04Z7vwnc,2,audio,6.0269999504089355,6.008905370493193,592,74,61,0.1,74,0.9971521496772766,ok,f5993ce3a6ee56298462c08692a96f78c7d914d7264257ed401de7d6e17fd148
|
||||||
|
-ri04Z7vwnc/2,-ri04Z7vwnc,2,vision,6.0269999504089355,2.106416702270508,30,35,61,0.1,35,0.9906440377235413,ok,f5993ce3a6ee56298462c08692a96f78c7d914d7264257ed401de7d6e17fd148
|
||||||
|
-ri04Z7vwnc/5,-ri04Z7vwnc,5,text,3.5450000762939453,1.8027290821075441,9,768,36,0.1,768,0.721091628074646,ok,2a5e4319bf592a18c8eb5117f7c7ec30c6c79528f9eac9fe4fb85bf70f277736
|
||||||
|
-ri04Z7vwnc/5,-ri04Z7vwnc,5,audio,3.5450000762939453,3.534527108637063,344,74,36,0.1,74,0.9974266886711121,ok,2a5e4319bf592a18c8eb5117f7c7ec30c6c79528f9eac9fe4fb85bf70f277736
|
||||||
|
-ri04Z7vwnc/5,-ri04Z7vwnc,5,vision,3.5450000762939453,3.5450000762939453,18,35,36,0.1,35,1.0,ok,2a5e4319bf592a18c8eb5117f7c7ec30c6c79528f9eac9fe4fb85bf70f277736
|
||||||
|
-s9qJ7ATP7w/1,-s9qJ7ATP7w,1,text,4.7919921875,2.610484106093645,11,768,48,0.1,768,0.7055363059043884,ok,a33e8f82a185bcb7648b523927308cdc2142e3076479ba2b45cd0e828636fc09
|
||||||
|
-s9qJ7ATP7w/1,-s9qJ7ATP7w,1,audio,4.7919921875,4.72247887628304,465,74,48,0.1,74,0.9860281348228455,ok,a33e8f82a185bcb7648b523927308cdc2142e3076479ba2b45cd0e828636fc09
|
||||||
|
-s9qJ7ATP7w/1,-s9qJ7ATP7w,1,vision,4.7919921875,4.7919921875,35,35,48,0.1,35,1.0,ok,a33e8f82a185bcb7648b523927308cdc2142e3076479ba2b45cd0e828636fc09
|
||||||
|
-s9qJ7ATP7w/0,-s9qJ7ATP7w,0,text,6.800000190734863,3.6206328462809334,24,768,69,0.1,768,0.6351987719535828,ok,f2324dbc6debe0d1e4254e6f943523373e3cbb30d8d0555413ce3a7b370a4af5
|
||||||
|
-s9qJ7ATP7w/0,-s9qJ7ATP7w,0,audio,6.800000190734863,6.715675868979983,672,74,69,0.1,74,0.9885490536689758,ok,f2324dbc6debe0d1e4254e6f943523373e3cbb30d8d0555413ce3a7b370a4af5
|
||||||
|
-s9qJ7ATP7w/0,-s9qJ7ATP7w,0,vision,6.800000190734863,6.199999928474425,56,35,69,0.1,35,0.984375,ok,f2324dbc6debe0d1e4254e6f943523373e3cbb30d8d0555413ce3a7b370a4af5
|
||||||
|
-s9qJ7ATP7w/5,-s9qJ7ATP7w,5,text,8.561002731323242,3.4511515218298876,19,768,86,0.1,768,0.5150972604751587,ok,bd35eb8de64ce1f8b27921294598f7cf5bb971255738f51d06458c7096c58eae
|
||||||
|
-s9qJ7ATP7w/5,-s9qJ7ATP7w,5,audio,8.561002731323242,8.46259727510246,840,74,86,0.1,74,0.9896790981292725,ok,bd35eb8de64ce1f8b27921294598f7cf5bb971255738f51d06458c7096c58eae
|
||||||
|
-s9qJ7ATP7w/5,-s9qJ7ATP7w,5,vision,8.561002731323242,8.561002731323242,72,35,86,0.1,35,1.0,ok,bd35eb8de64ce1f8b27921294598f7cf5bb971255738f51d06458c7096c58eae
|
||||||
|
-s9qJ7ATP7w/4,-s9qJ7ATP7w,4,text,4.427018165588379,2.2374289706349377,9,768,45,0.1,768,0.6215080618858337,ok,76e0672ffce5857a789b41fb10606402f7da58c33579b1032a8282f86fb35bc9
|
||||||
|
-s9qJ7ATP7w/4,-s9qJ7ATP7w,4,audio,4.427018165588379,4.367301462550421,425,74,45,0.1,74,0.9892621040344238,ok,76e0672ffce5857a789b41fb10606402f7da58c33579b1032a8282f86fb35bc9
|
||||||
|
-s9qJ7ATP7w/4,-s9qJ7ATP7w,4,vision,4.427018165588379,3.027018189430237,31,35,45,0.1,35,0.96875,ok,76e0672ffce5857a789b41fb10606402f7da58c33579b1032a8282f86fb35bc9
|
||||||
|
-s9qJ7ATP7w/7,-s9qJ7ATP7w,7,text,6.966015815734863,3.9709798723459264,28,768,70,0.1,768,0.661829948425293,ok,5301f4403cf7cd299cc525d2f443392ac70fd70c3c9cd4a09406f2ba4df6734d
|
||||||
|
-s9qJ7ATP7w/7,-s9qJ7ATP7w,7,audio,6.966015815734863,6.848488390969263,684,74,70,0.1,74,0.9849807620048523,ok,5301f4403cf7cd299cc525d2f443392ac70fd70c3c9cd4a09406f2ba4df6734d
|
||||||
|
-s9qJ7ATP7w/7,-s9qJ7ATP7w,7,vision,6.966015815734863,4.0493486523628235,57,35,70,0.1,35,0.9496123194694519,ok,5301f4403cf7cd299cc525d2f443392ac70fd70c3c9cd4a09406f2ba4df6734d
|
||||||
|
-s9qJ7ATP7w/6,-s9qJ7ATP7w,6,text,2.4749999046325684,1.6227161765098577,5,768,25,0.1,768,0.8113580942153931,ok,40b684fd1b4559f0f82e72f9b47f4ce7c5e15059fe8d0d8ab112637cd3386451
|
||||||
|
-s9qJ7ATP7w/6,-s9qJ7ATP7w,6,audio,2.4749999046325684,2.4289188214250514,232,74,25,0.1,74,0.9844902753829956,ok,40b684fd1b4559f0f82e72f9b47f4ce7c5e15059fe8d0d8ab112637cd3386451
|
||||||
|
-s9qJ7ATP7w/6,-s9qJ7ATP7w,6,vision,2.4749999046325684,1.8999999761581425,14,35,25,0.1,35,1.0,ok,40b684fd1b4559f0f82e72f9b47f4ce7c5e15059fe8d0d8ab112637cd3386451
|
||||||
|
-s9qJ7ATP7w/8,-s9qJ7ATP7w,8,text,3.2949869632720947,1.9395141303539274,13,768,33,0.1,768,0.6465047001838684,ok,1de4d91bd9d9fb508a3779658747da88460db0bdb1beefe7ea5b2c77edfd2239
|
||||||
|
-s9qJ7ATP7w/8,-s9qJ7ATP7w,8,audio,3.2949869632720947,3.253163003921509,319,74,33,0.1,74,0.9880942702293396,ok,1de4d91bd9d9fb508a3779658747da88460db0bdb1beefe7ea5b2c77edfd2239
|
||||||
|
-s9qJ7ATP7w/8,-s9qJ7ATP7w,8,vision,3.2949869632720947,3.016666531562805,20,35,33,0.1,35,0.9731181859970093,ok,1de4d91bd9d9fb508a3779658747da88460db0bdb1beefe7ea5b2c77edfd2239
|
||||||
|
-yRb-Jum7EQ/1,-yRb-Jum7EQ,1,text,29.288021087646484,15.053765958920112,49,768,293,0.1,768,0.6873865723609924,ok,2f16da1baa8660e5bfc608f5237cdd59725d16e57a19bc02c4556cd2c7129a46
|
||||||
|
-yRb-Jum7EQ/1,-yRb-Jum7EQ,1,audio,29.288021087646484,28.759709144524624,2911,74,293,0.1,74,0.9836021065711975,ok,2f16da1baa8660e5bfc608f5237cdd59725d16e57a19bc02c4556cd2c7129a46
|
||||||
|
-yRb-Jum7EQ/1,-yRb-Jum7EQ,1,vision,29.288021087646484,27.471354246139526,178,35,293,0.1,35,0.999393880367279,ok,2f16da1baa8660e5bfc608f5237cdd59725d16e57a19bc02c4556cd2c7129a46
|
||||||
|
-yRb-Jum7EQ/5,-yRb-Jum7EQ,5,text,13.169010162353516,8.127011682093151,25,768,132,0.1,768,0.7192046046257019,ok,6d6145293cf84d1fc8324dfa8d39465e2d7fc07affb67ff68deeabf6e3c410d9
|
||||||
|
-yRb-Jum7EQ/5,-yRb-Jum7EQ,5,audio,13.169010162353516,12.971401820754682,1305,74,132,0.1,74,0.9867846369743347,ok,6d6145293cf84d1fc8324dfa8d39465e2d7fc07affb67ff68deeabf6e3c410d9
|
||||||
|
-yRb-Jum7EQ/5,-yRb-Jum7EQ,5,vision,13.169010162353516,13.169010162353516,97,35,132,0.1,35,1.0,ok,6d6145293cf84d1fc8324dfa8d39465e2d7fc07affb67ff68deeabf6e3c410d9
|
||||||
|
-yRb-Jum7EQ/6,-yRb-Jum7EQ,6,text,11.241994857788086,4.740236129239202,19,768,113,0.1,768,0.6869907975196838,ok,b51eb9ad06de804934cb6a315de591ba9ff7f8aa71bffff7f8001ab29945d989
|
||||||
|
-yRb-Jum7EQ/6,-yRb-Jum7EQ,6,audio,11.241994857788086,11.070711010679133,1125,74,113,0.1,74,0.9854583740234375,ok,b51eb9ad06de804934cb6a315de591ba9ff7f8aa71bffff7f8001ab29945d989
|
||||||
|
-yRb-Jum7EQ/6,-yRb-Jum7EQ,6,vision,11.241994857788086,11.241994857788086,87,35,113,0.1,35,1.0,ok,b51eb9ad06de804934cb6a315de591ba9ff7f8aa71bffff7f8001ab29945d989
|
||||||
|
@@ -0,0 +1,101 @@
|
|||||||
|
sample_id,video_id,clip_id,true_polarity,true_polarity_name,true_sentiment,B0_predicted_polarity,B1_predicted_polarity,B2_predicted_polarity,B3_predicted_polarity,B4_predicted_polarity,B0_predicted_sentiment,B1_predicted_sentiment,B2_predicted_sentiment,B3_predicted_sentiment,B4_predicted_sentiment
|
||||||
|
-3g5yACwYnA/13,-3g5yACwYnA,13,2,positive,0.6666666865348816,2,2,2,2,2,0.13452544808387756,0.16021420061588287,0.14872726798057556,0.16442018747329712,0.229058176279068
|
||||||
|
-3g5yACwYnA/3,-3g5yACwYnA,3,1,neutral,0.0,2,2,2,2,2,0.3026959300041199,0.33461102843284607,0.42589470744132996,0.31887251138687134,0.29829922318458557
|
||||||
|
-3g5yACwYnA/2,-3g5yACwYnA,2,1,neutral,0.0,2,2,2,2,2,0.3966212272644043,0.46180492639541626,0.48163092136383057,0.43820273876190186,0.5120640397071838
|
||||||
|
-3g5yACwYnA/9,-3g5yACwYnA,9,2,positive,0.6666666865348816,2,2,2,2,2,0.6833762526512146,0.7414957880973816,0.7930513024330139,0.7018600106239319,0.6162992715835571
|
||||||
|
-3nNcZdcdvU/5,-3nNcZdcdvU,5,1,neutral,0.0,0,0,0,0,0,-0.7614330053329468,-0.6502371430397034,-0.6979431509971619,-0.6381849050521851,-0.5865846872329712
|
||||||
|
-HwX2H8Z4hY/2,-HwX2H8Z4hY,2,1,neutral,0.0,2,2,2,2,2,0.6054284572601318,0.5735942721366882,0.6148816347122192,0.5838509798049927,0.8063328266143799
|
||||||
|
-HwX2H8Z4hY/5,-HwX2H8Z4hY,5,0,negative,-1.0,0,0,2,2,2,0.08644700050354004,-0.01681867241859436,0.054207056760787964,0.008860379457473755,0.08198326826095581
|
||||||
|
-HwX2H8Z4hY/6,-HwX2H8Z4hY,6,1,neutral,0.0,2,2,2,2,2,0.31941771507263184,0.4352392554283142,0.45426928997039795,0.40970659255981445,0.5301598310470581
|
||||||
|
-HwX2H8Z4hY/9,-HwX2H8Z4hY,9,0,negative,-0.3333333432674408,0,0,0,0,0,-0.07338133454322815,-0.14210763573646545,-0.13199511170387268,-0.13720062375068665,-0.2480643391609192
|
||||||
|
-NFrJFQijFE/1,-NFrJFQijFE,1,2,positive,0.3333333432674408,2,2,2,2,1,-0.029091298580169678,0.04221409559249878,0.023884594440460205,0.038062989711761475,-0.10948124527931213
|
||||||
|
-NFrJFQijFE/2,-NFrJFQijFE,2,1,neutral,0.0,2,2,2,2,2,0.5534588098526001,0.6096377372741699,0.5728316307067871,0.5841389894485474,0.6658575534820557
|
||||||
|
-THoVjtIkeU/12,-THoVjtIkeU,12,2,positive,2.0,2,2,2,2,2,0.1105184555053711,-0.015041038393974304,0.01798933744430542,0.013076543807983398,0.05507504940032959
|
||||||
|
-THoVjtIkeU/2,-THoVjtIkeU,2,2,positive,1.3333333730697632,2,2,2,2,2,0.14952212572097778,-0.03578713536262512,0.21528872847557068,-0.047229185700416565,0.07181030511856079
|
||||||
|
-THoVjtIkeU/6,-THoVjtIkeU,6,2,positive,2.6666667461395264,2,2,2,2,2,0.28439241647720337,0.10436569154262543,0.19625744223594666,0.109825998544693,0.11878173053264618
|
||||||
|
-UuX1xuaiiE/1,-UuX1xuaiiE,1,2,positive,1.6666666269302368,2,2,2,2,2,0.05641767382621765,0.05484718084335327,0.02277640998363495,0.05113929510116577,0.011383742094039917
|
||||||
|
-UuX1xuaiiE/0,-UuX1xuaiiE,0,2,positive,2.6666667461395264,2,2,2,2,2,0.6992747783660889,0.6104735136032104,0.7187632918357849,0.6213169097900391,0.5144999027252197
|
||||||
|
-UuX1xuaiiE/3,-UuX1xuaiiE,3,2,positive,0.3333333432674408,2,2,2,2,2,0.66947340965271,0.5266174077987671,0.5624967217445374,0.5222238898277283,0.45163899660110474
|
||||||
|
-UuX1xuaiiE/6,-UuX1xuaiiE,6,0,negative,-0.3333333432674408,2,2,2,2,2,0.12534436583518982,0.002005934715270996,-0.03466030955314636,-0.01757238805294037,0.03905734419822693
|
||||||
|
-a55Q6RWvTA/3,-a55Q6RWvTA,3,2,positive,0.6666666865348816,2,2,2,2,2,0.24422422051429749,0.13229034841060638,0.1552799940109253,0.1265096515417099,0.17101362347602844
|
||||||
|
-aNfi7CP8vM/7,-aNfi7CP8vM,7,1,neutral,0.0,2,2,2,2,2,0.7103666663169861,0.8981346487998962,0.7498640418052673,0.908927321434021,0.7986767292022705
|
||||||
|
-aqamKhZ1Ec/0,-aqamKhZ1Ec,0,0,negative,-2.0,2,2,2,2,2,0.14777150750160217,0.2721712589263916,0.20749390125274658,0.26992136240005493,0.2557317018508911
|
||||||
|
-dxfTGcXJoc/1,-dxfTGcXJoc,1,2,positive,1.6666666269302368,2,2,2,2,2,0.4534122347831726,0.15173421800136566,0.23860448598861694,0.08927315473556519,0.17231377959251404
|
||||||
|
-dxfTGcXJoc/0,-dxfTGcXJoc,0,2,positive,1.0,2,2,2,1,2,0.30912163853645325,0.18118944764137268,0.1967533975839615,0.17624947428703308,0.20736707746982574
|
||||||
|
-dxfTGcXJoc/2,-dxfTGcXJoc,2,2,positive,0.6666666865348816,2,2,2,2,2,0.5904080867767334,0.34547320008277893,0.4100586771965027,0.3144497573375702,0.33179861307144165
|
||||||
|
-dxfTGcXJoc/6,-dxfTGcXJoc,6,2,positive,0.6666666865348816,2,2,2,2,2,0.40664565563201904,0.07521255314350128,0.0900307297706604,0.107652947306633,0.09203940629959106
|
||||||
|
-egA8-b7-3M/26,-egA8-b7-3M,26,2,positive,1.6666666269302368,2,2,2,2,2,0.8926460146903992,0.8544832468032837,0.7031165957450867,0.8596105575561523,0.6629092693328857
|
||||||
|
-egA8-b7-3M/17,-egA8-b7-3M,17,2,positive,0.3333333432674408,2,2,2,2,2,0.780834972858429,0.8010746240615845,0.7224219441413879,0.7491263151168823,0.7902140617370605
|
||||||
|
-egA8-b7-3M/18,-egA8-b7-3M,18,2,positive,1.0,2,2,2,2,2,0.8284801244735718,0.8428431153297424,0.7629272937774658,0.8139240741729736,0.759922981262207
|
||||||
|
-egA8-b7-3M/16,-egA8-b7-3M,16,2,positive,0.6666666865348816,2,2,2,2,2,1.1456997394561768,1.103100299835205,0.9217796921730042,1.0922913551330566,0.9465919733047485
|
||||||
|
-egA8-b7-3M/13,-egA8-b7-3M,13,2,positive,0.3333333432674408,2,2,2,2,2,0.29149097204208374,0.2857578694820404,0.29504942893981934,0.2710241973400116,0.433204710483551
|
||||||
|
-egA8-b7-3M/1,-egA8-b7-3M,1,2,positive,0.6666666865348816,2,2,2,2,2,0.49869486689567566,0.5516833662986755,0.4600910544395447,0.5341068506240845,0.5454411506652832
|
||||||
|
-egA8-b7-3M/6,-egA8-b7-3M,6,2,positive,0.6666666865348816,2,2,2,2,2,0.8187044262886047,0.8529066443443298,0.7433335781097412,0.8614661693572998,0.7032537460327148
|
||||||
|
-egA8-b7-3M/9,-egA8-b7-3M,9,2,positive,0.3333333432674408,2,2,2,2,2,0.746141791343689,0.755176305770874,0.7585423588752747,0.7326465845108032,0.7642356157302856
|
||||||
|
-egA8-b7-3M/20,-egA8-b7-3M,20,2,positive,0.3333333432674408,2,2,2,2,2,0.75407874584198,0.783722460269928,0.7880091667175293,0.7994571924209595,0.7963440418243408
|
||||||
|
-iRBcNs9oI8/3,-iRBcNs9oI8,3,2,positive,0.6666666865348816,2,2,2,2,2,0.41121792793273926,0.26391637325286865,0.16125774383544922,0.30018502473831177,0.22292546927928925
|
||||||
|
-iRBcNs9oI8/7,-iRBcNs9oI8,7,2,positive,0.6666666865348816,2,2,2,2,2,0.6934136152267456,0.5427778959274292,0.5866559743881226,0.5839468240737915,0.5824995636940002
|
||||||
|
-iRBcNs9oI8/6,-iRBcNs9oI8,6,2,positive,0.3333333432674408,2,2,2,2,2,0.4273873567581177,0.3189065456390381,0.3010154962539673,0.3516676127910614,0.32149577140808105
|
||||||
|
-iRBcNs9oI8/9,-iRBcNs9oI8,9,2,positive,1.3333333730697632,2,2,1,2,2,0.01312372088432312,-0.0918610543012619,0.1977950930595398,-0.06410297751426697,-0.14683973789215088
|
||||||
|
-iRBcNs9oI8/8,-iRBcNs9oI8,8,2,positive,1.0,2,2,2,2,2,0.45394086837768555,0.2737142741680145,0.34107109904289246,0.34081369638442993,0.36690521240234375
|
||||||
|
-lzEya4AM_4/5,-lzEya4AM_4,5,0,negative,-0.6666666865348816,2,1,1,1,2,0.003696233034133911,0.07954774796962738,-0.02835392951965332,0.06377251446247101,0.0930773913860321
|
||||||
|
-lzEya4AM_4/6,-lzEya4AM_4,6,0,negative,-0.3333333432674408,1,1,1,1,1,-0.08578968048095703,-0.05591443181037903,-0.16103187203407288,-0.07273416221141815,-0.03954911231994629
|
||||||
|
-mJ2ud6oKI8/1,-mJ2ud6oKI8,1,1,neutral,0.0,2,2,2,2,1,0.32054486870765686,0.31638309359550476,0.4096267521381378,0.32599377632141113,0.11944134533405304
|
||||||
|
-mJ2ud6oKI8/2,-mJ2ud6oKI8,2,1,neutral,0.0,2,2,2,2,1,0.3205912411212921,0.2857860326766968,0.32545098662376404,0.2961588501930237,0.07219411432743073
|
||||||
|
-mJ2ud6oKI8/6,-mJ2ud6oKI8,6,2,positive,0.3333333432674408,2,2,2,2,2,0.34881988167762756,0.2555202841758728,0.19199120998382568,0.22443163394927979,0.11697705090045929
|
||||||
|
-mJ2ud6oKI8/9,-mJ2ud6oKI8,9,1,neutral,0.0,2,2,2,2,2,0.49129533767700195,0.509657621383667,0.40949302911758423,0.49832504987716675,0.3923254609107971
|
||||||
|
-mJ2ud6oKI8/8,-mJ2ud6oKI8,8,0,negative,-0.3333333432674408,0,0,0,0,0,0.033514365553855896,-0.12224322557449341,0.015729933977127075,-0.13064196705818176,-0.10159948468208313
|
||||||
|
-mqbVkbCndg/0,-mqbVkbCndg,0,1,neutral,0.0,2,2,2,2,2,-0.09232228994369507,-0.27456536889076233,-0.23857474327087402,-0.2712118625640869,-0.06543052196502686
|
||||||
|
-t217m2on-s/2,-t217m2on-s,2,2,positive,0.6666666865348816,2,2,2,2,2,0.4828706979751587,0.4535056948661804,0.48303860425949097,0.4338659644126892,0.29441648721694946
|
||||||
|
-t217m2on-s/7,-t217m2on-s,7,0,negative,-0.6666666865348816,2,2,2,2,2,0.03153911232948303,0.12483498454093933,0.23645390570163727,0.14657872915267944,0.2676320970058441
|
||||||
|
-tANM6ETl_M/3,-tANM6ETl_M,3,2,positive,0.3333333432674408,2,2,2,2,2,0.13329337537288666,0.27301740646362305,0.20109084248542786,0.23898470401763916,0.2801833152770996
|
||||||
|
-tPCytz4rww/11,-tPCytz4rww,11,1,neutral,0.0,2,2,2,2,2,0.39490315318107605,0.41642117500305176,0.5103898644447327,0.40319329500198364,0.5221617221832275
|
||||||
|
-tPCytz4rww/10,-tPCytz4rww,10,1,neutral,0.0,2,2,2,2,2,0.45396602153778076,0.48757827281951904,0.39305949211120605,0.48652955889701843,0.5163553357124329
|
||||||
|
-tPCytz4rww/12,-tPCytz4rww,12,1,neutral,0.0,2,2,2,2,2,0.4888620376586914,0.47673290967941284,0.523354709148407,0.47867047786712646,0.5164443254470825
|
||||||
|
-tPCytz4rww/16,-tPCytz4rww,16,1,neutral,0.0,2,2,2,2,2,0.16386589407920837,0.1889176368713379,0.15202002227306366,0.20772001147270203,0.10384060442447662
|
||||||
|
-tPCytz4rww/18,-tPCytz4rww,18,1,neutral,0.0,2,2,2,2,2,0.39335885643959045,0.48845911026000977,0.4801775813102722,0.49252307415008545,0.5596518516540527
|
||||||
|
-vxjVxOeScU/4,-vxjVxOeScU,4,2,positive,1.0,2,2,2,2,2,1.300539255142212,1.3340058326721191,1.3243328332901,1.3665581941604614,1.3101918697357178
|
||||||
|
-wMB_hJL-3o/7,-wMB_hJL-3o,7,2,positive,0.3333333432674408,2,2,2,2,2,0.43083706498146057,0.3697093427181244,0.3281114399433136,0.3639947175979614,0.4309353828430176
|
||||||
|
-wny0OAz3g8/1,-wny0OAz3g8,1,2,positive,0.3333333432674408,2,2,2,2,2,-0.029849499464035034,-0.0726122260093689,-0.024748489260673523,-0.07751013338565826,0.01515229046344757
|
||||||
|
-wny0OAz3g8/0,-wny0OAz3g8,0,2,positive,0.3333333432674408,1,1,1,1,1,-0.26475363969802856,-0.4248991906642914,-0.3667675256729126,-0.43525469303131104,-0.2892696261405945
|
||||||
|
-wny0OAz3g8/3,-wny0OAz3g8,3,0,negative,-1.3333333730697632,2,2,2,2,2,-0.03229987621307373,0.011097535490989685,-0.044357895851135254,0.014330193400382996,0.1011495590209961
|
||||||
|
-wny0OAz3g8/2,-wny0OAz3g8,2,0,negative,-0.3333333432674408,2,2,2,2,2,0.0680823028087616,0.020685836672782898,0.1548982560634613,0.049939438700675964,0.12679153680801392
|
||||||
|
-wny0OAz3g8/5,-wny0OAz3g8,5,2,positive,1.0,2,2,2,2,2,0.4779983460903168,0.42994704842567444,0.3525465130805969,0.4249609410762787,0.4561484158039093
|
||||||
|
-wny0OAz3g8/7,-wny0OAz3g8,7,1,neutral,0.0,2,2,2,2,2,0.14757554233074188,0.11216723173856735,0.15002727508544922,0.10439274460077286,0.08023504912853241
|
||||||
|
-wny0OAz3g8/9,-wny0OAz3g8,9,2,positive,0.6666666865348816,2,2,2,2,2,0.18945443630218506,0.2243257611989975,0.255195677280426,0.2078581154346466,0.26747772097587585
|
||||||
|
-571d8cVauQ/0,-571d8cVauQ,0,1,neutral,0.0,1,1,1,1,1,0.1813371479511261,0.26319703459739685,0.20845845341682434,0.2597145736217499,0.2923925817012787
|
||||||
|
-571d8cVauQ/5,-571d8cVauQ,5,2,positive,0.3333333432674408,2,2,2,2,2,0.7947983145713806,0.8263474702835083,0.8470317125320435,0.8056427836418152,0.7695794105529785
|
||||||
|
-I_e4mIh0yE/1,-I_e4mIh0yE,1,0,negative,-0.6666666865348816,1,1,2,1,2,-0.053890615701675415,0.05120442807674408,0.14258822798728943,0.058895498514175415,-0.0020049959421157837
|
||||||
|
-I_e4mIh0yE/3,-I_e4mIh0yE,3,2,positive,0.6666666865348816,2,2,2,2,2,0.20941728353500366,0.2584538459777832,0.3930857181549072,0.27197059988975525,0.3383082151412964
|
||||||
|
-UacrmKiTn4/10,-UacrmKiTn4,10,0,negative,-0.6666666865348816,2,2,2,2,2,-0.007775545120239258,0.10976105183362961,0.2304498553276062,0.1546456217765808,0.1571420133113861
|
||||||
|
-UacrmKiTn4/4,-UacrmKiTn4,4,0,negative,-1.3333333730697632,2,2,2,2,2,0.24989673495292664,0.44462817907333374,0.4752929210662842,0.4392329156398773,0.5184646844863892
|
||||||
|
-hnBHBN8p5A/7,-hnBHBN8p5A,7,2,positive,0.6666666865348816,2,2,2,2,2,0.25469517707824707,0.267363965511322,0.26611560583114624,0.2937210500240326,0.4176241457462311
|
||||||
|
-hnBHBN8p5A/6,-hnBHBN8p5A,6,2,positive,1.0,2,2,2,2,2,0.7116098403930664,0.6368920803070068,0.8011423349380493,0.6396559476852417,0.7429358959197998
|
||||||
|
-qDkUB0GgYY/6,-qDkUB0GgYY,6,2,positive,1.3333333730697632,2,2,2,2,2,1.4230821132659912,1.456520915031433,1.482230544090271,1.4008221626281738,1.3085699081420898
|
||||||
|
-uywlfIYOS8/4,-uywlfIYOS8,4,2,positive,0.3333333432674408,2,2,2,2,2,0.6702563762664795,0.7197655439376831,0.6297314763069153,0.7727823257446289,0.7619336843490601
|
||||||
|
-6rXp3zJ3kc/8,-6rXp3zJ3kc,8,0,negative,-1.0,2,2,2,2,2,0.23289291560649872,0.10650955140590668,-0.048172906041145325,0.1659325510263443,0.07517096400260925
|
||||||
|
-9y-fZ3swSY/0,-9y-fZ3swSY,0,2,positive,1.0,2,2,2,2,2,0.8935344815254211,0.981560468673706,1.0655876398086548,0.9575732946395874,0.960616946220398
|
||||||
|
-9y-fZ3swSY/4,-9y-fZ3swSY,4,1,neutral,0.0,2,2,2,2,1,0.7359274625778198,0.8043063282966614,0.8141930103302002,0.8211376667022705,0.801345944404602
|
||||||
|
-9y-fZ3swSY/8,-9y-fZ3swSY,8,2,positive,1.6666666269302368,2,2,2,2,2,0.0990075170993805,0.17485348880290985,0.26351824402809143,0.15360157191753387,0.2083582729101181
|
||||||
|
-AUZQgSxyPQ/2,-AUZQgSxyPQ,2,2,positive,0.6666666865348816,2,2,2,2,2,0.9318890571594238,1.1048165559768677,1.1751078367233276,1.091732382774353,0.7978025078773499
|
||||||
|
-HeZS2-Prhc/2,-HeZS2-Prhc,2,2,positive,0.6666666865348816,2,2,2,2,2,0.42690128087997437,0.37649014592170715,0.3945169448852539,0.3838789463043213,0.30630016326904297
|
||||||
|
-MeTTeMJBNc/0,-MeTTeMJBNc,0,2,positive,0.3333333432674408,2,2,2,2,2,0.5873996019363403,0.4888506829738617,0.520802915096283,0.4796018600463867,0.48370563983917236
|
||||||
|
-MeTTeMJBNc/13,-MeTTeMJBNc,13,2,positive,0.3333333432674408,2,2,2,2,2,0.681941032409668,0.7230368852615356,0.6091403365135193,0.7292149662971497,0.745182454586029
|
||||||
|
-MeTTeMJBNc/7,-MeTTeMJBNc,7,2,positive,1.3333333730697632,2,2,2,2,2,0.5029646158218384,0.47240275144577026,0.5074828863143921,0.46745914220809937,0.47605717182159424
|
||||||
|
-RfYyzHpjk4/11,-RfYyzHpjk4,11,2,positive,2.0,2,2,2,2,2,0.5564365983009338,0.42342162132263184,0.5021690726280212,0.43483608961105347,0.43543678522109985
|
||||||
|
-RfYyzHpjk4/8,-RfYyzHpjk4,8,2,positive,0.3333333432674408,2,2,2,2,2,0.3994223177433014,0.3921676576137543,0.2823079228401184,0.40095123648643494,0.42251622676849365
|
||||||
|
-RfYyzHpjk4/2,-RfYyzHpjk4,2,2,positive,0.3333333432674408,2,2,2,2,1,0.365897536277771,0.38750067353248596,0.23726055026054382,0.3845585286617279,0.21417461335659027
|
||||||
|
-UUCSKoHeMA/0,-UUCSKoHeMA,0,1,neutral,0.0,2,2,2,2,2,0.5638669729232788,0.5825111269950867,0.6482444405555725,0.5764590501785278,0.5579362511634827
|
||||||
|
-ri04Z7vwnc/0,-ri04Z7vwnc,0,1,neutral,0.0,2,2,2,2,1,0.26682645082473755,0.1752839833498001,0.11657405644655228,0.17440180480480194,-0.12501448392868042
|
||||||
|
-ri04Z7vwnc/2,-ri04Z7vwnc,2,0,negative,-0.3333333432674408,2,2,2,2,1,0.41648173332214355,0.39698678255081177,0.31453409790992737,0.42282843589782715,-0.06992921233177185
|
||||||
|
-ri04Z7vwnc/5,-ri04Z7vwnc,5,2,positive,0.3333333432674408,2,2,2,2,2,0.7894954681396484,0.9298861026763916,0.8353837132453918,0.9478996992111206,0.9307202100753784
|
||||||
|
-s9qJ7ATP7w/1,-s9qJ7ATP7w,1,2,positive,0.6666666865348816,2,2,2,2,2,0.8451777696609497,0.7406796216964722,0.6926915645599365,0.7564423084259033,0.7351438403129578
|
||||||
|
-s9qJ7ATP7w/0,-s9qJ7ATP7w,0,0,negative,-0.3333333432674408,2,2,2,2,0,0.08492019772529602,0.08116327226161957,0.09997616708278656,0.09164288640022278,-0.18887805938720703
|
||||||
|
-s9qJ7ATP7w/5,-s9qJ7ATP7w,5,0,negative,-2.0,2,0,2,0,2,0.25190818309783936,0.23056381940841675,0.42701321840286255,0.2685755491256714,0.35108593106269836
|
||||||
|
-s9qJ7ATP7w/4,-s9qJ7ATP7w,4,1,neutral,0.0,2,2,2,2,2,0.8093756437301636,0.6910800933837891,0.8017246723175049,0.6906447410583496,0.7730451822280884
|
||||||
|
-s9qJ7ATP7w/7,-s9qJ7ATP7w,7,2,positive,1.0,2,2,2,2,1,0.04829922318458557,-0.06272962689399719,0.04885163903236389,-0.056273818016052246,-0.3453366756439209
|
||||||
|
-s9qJ7ATP7w/6,-s9qJ7ATP7w,6,1,neutral,0.0,2,2,2,2,2,-0.11959794163703918,0.12442992627620697,0.22614140808582306,0.12353259325027466,0.05900001525878906
|
||||||
|
-s9qJ7ATP7w/8,-s9qJ7ATP7w,8,2,positive,1.3333333730697632,2,2,2,2,1,0.5094736218452454,0.4484489858150482,0.6055055856704712,0.42063793540000916,0.0980398952960968
|
||||||
|
-yRb-Jum7EQ/1,-yRb-Jum7EQ,1,1,neutral,0.0,0,0,0,0,0,-0.39470893144607544,-0.5162537693977356,-0.5732794404029846,-0.5318436622619629,-0.47799739241600037
|
||||||
|
-yRb-Jum7EQ/5,-yRb-Jum7EQ,5,0,negative,-0.6666666865348816,1,1,1,1,1,-0.31519562005996704,-0.4650754928588867,-0.48384714126586914,-0.4728321433067322,-0.4012225568294525
|
||||||
|
-yRb-Jum7EQ/6,-yRb-Jum7EQ,6,1,neutral,0.0,0,0,0,0,0,-0.3192596435546875,-0.48287540674209595,-0.5565515160560608,-0.4418525695800781,-0.5206273794174194
|
||||||
|
@@ -0,0 +1,61 @@
|
|||||||
|
# 问题一 B0-B4 模型对比结果
|
||||||
|
|
||||||
|
本报告按 `math/问题一.pdf` 第 4.8.3 节实施五组单因素对照。全部 100 条附件一视频均使用同一套冻结特征抽取器、0.1 秒主时间网格、观测掩码与按 `video_id` 分组的五折划分。情感标签仅用于折内的轻量预测探针,不进入时间对齐、CTC 后验或动态签名计算。
|
||||||
|
|
||||||
|
## 实际特征和比较版本
|
||||||
|
|
||||||
|
| 版本 | 实施内容 |
|
||||||
|
| --- | --- |
|
||||||
|
| B0 | CTC Viterbi 词区间与音频/视觉物理时间区间投影;训练折均值/标准差缩放;视觉旋转向量和视线向量作普通分量均值。 |
|
||||||
|
| B1 | 在 B0 上改用训练折中位数与 MAD 的稳健尺度(MAD 为零时回退到训练折标准差),并对视觉旋转用 SO(3) 均值、视线用归一化向量均值。 |
|
||||||
|
| B2 | 在 B1 上只将文本词向量的硬边界投影替换为固定转写 CTC 状态图的前向–后向占据概率投影;媒体时间戳仍按物理时间。 |
|
||||||
|
| B3 | 在 B1 上仅附加音频 log-F0/能量与视觉 jaw-open/嘴部开合率的时间增广一、二阶路径签名;缺测点之间不连线。 |
|
||||||
|
| B4 | 在 B1 上仅附加冻结 Wav2Vec2-base-960h 最后四层的时间投影表示。 |
|
||||||
|
|
||||||
|
附件一现有标签表和原始视频被直接使用。文本为 768 维 BERT-base-uncased 末四层均值;音频为 74 维(log-Mel 40、MFCC 13、ΔMFCC 13、韵律/谱统计 8);视觉为 35 维(17 个 MediaPipe blendshape 代理、6 维头姿、6 维近似视线、6 维面部比例几何)。视觉代理和几何索引定义见 `compare_models.py`,它们与 OpenFace AU 定义并不等同;该限制须在论文中明示。
|
||||||
|
|
||||||
|
## 结果
|
||||||
|
|
||||||
|
分类按连续标签的严格符号构造 Negative/Neutral/Positive,强度预测限幅到 [-3, 3]。每折训练折单独拟合标准化器、逻辑回归(C=0.05)与 Ridge(alpha=25);固定五折由 GroupKFold 按原始 `video_id` 分组。总体 OOF 指标按 100 条留组预测计算。
|
||||||
|
|
||||||
|
| 方法 | OOF Accuracy | OOF Macro-F1 | OOF MAE | OOF Pearson | Macro-F1 折均值±SD |
|
||||||
|
| --- | ---: | ---: | ---: | ---: | ---: |
|
||||||
|
| B0 | 0.600 | 0.361 | 0.564 | 0.346 | 0.337 ± 0.098 |
|
||||||
|
| B1 | 0.610 | 0.387 | 0.599 | 0.238 | 0.349 ± 0.119 |
|
||||||
|
| B2 | 0.580 | 0.331 | 0.588 | 0.279 | 0.313 ± 0.087 |
|
||||||
|
| B3 | 0.590 | 0.359 | 0.600 | 0.230 | 0.333 ± 0.096 |
|
||||||
|
| B4 | 0.600 | 0.422 | 0.593 | 0.199 | 0.398 ± 0.089 |
|
||||||
|
|
||||||
|
按当前 OOF 探针,Macro-F1 最高的是 B4(0.422),MAE 最低的是 B0(0.564)。相对 B1,B4 的 Macro-F1 差为 +0.035,视频组 Bootstrap 95% 区间 [-0.051, +0.137],区间跨过零。B0 的 MAE 差为 -0.035,区间 [-0.063, -0.012],Pearson 差为 +0.108。B2 的 Macro-F1 差为 -0.057,区间 [-0.109, -0.001],MAE 差为 -0.011;B3 的 Macro-F1 差为 -0.029,区间 [-0.072, +0.000]。组 Bootstrap 只反映当前 37 个来源视频上的抽样不确定性,未作多重比较校正。
|
||||||
|
|
||||||
|
## 覆盖与对齐核验
|
||||||
|
|
||||||
|
样本数:100;标签表连续标签与极性一致 100/100;CTC 硬对齐完整的样本 91/100。B0 硬区间平均文本覆盖率为 0.673,B2 平均词占据后验质量为 0.229(这是期望占据率,不是覆盖率);视觉检测覆盖率(按 0.1 秒网格平均)为 0.859。后验内部 90% 边界区间平均宽度为起点 0.131 秒、终点 0.111 秒。没有人工边界子集,因此这些宽度只是模型内部不确定性摘要,不是经验校准率或边界误差。
|
||||||
|
|
||||||
|
B0–B4 的特征来源、有效掩码、连续覆盖率、词边界后验宽度、折内预测及视频哈希均分别保存在配套 CSV/JSON 中。没有人工词边界时不报告 IoU/MATE,也不把不同方法的预测探针分数解释为对齐真值。
|
||||||
|
|
||||||
|
## 文件
|
||||||
|
|
||||||
|
- `comparison_summary.csv`:五种方法总体 OOF 与折均值指标。
|
||||||
|
- `fold_metrics.csv`:每折分类/回归指标与特征维数。
|
||||||
|
- `oof_predictions.csv`:每条样本的留组预测。
|
||||||
|
- `group_bootstrap_deltas.csv`:相对 B1 的视频组配对 Bootstrap 95% 区间。
|
||||||
|
- `sample_alignment_summary.csv`:100 条样本、硬/概率文本覆盖和后验边界宽度。
|
||||||
|
- `modality_summary.csv`:300 行样本-模态明细。
|
||||||
|
- `word_alignment_posterior.csv`:逐词硬边界、实际相对质量权重及后验边界区间。
|
||||||
|
- `split_assignments.csv`:逐样本折号。
|
||||||
|
- `comparison.png`:核心 OOF 指标图。
|
||||||
|
- `typical_alignment_example.png`:中位时长样本的词边界、覆盖率、声学/视觉轨迹与视频帧核验图。
|
||||||
|
- `run_manifest.json`:环境、模型 revision、参数、哈希与运行时间。
|
||||||
|
|
||||||
|
## 复现
|
||||||
|
|
||||||
|
在仓库根目录执行:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd math
|
||||||
|
uv sync
|
||||||
|
uv run python compare_models.py
|
||||||
|
```
|
||||||
|
|
||||||
|
特征缓存写在 `math/cache/native/`,正式结果只写在 `math/results/model_comparison/`。删除缓存后会从附件一重新提取。首次运行需要下载 BERT、Wav2Vec2 和 MediaPipe Face Landmarker 权重。
|
||||||
@@ -0,0 +1,293 @@
|
|||||||
|
{
|
||||||
|
"created_at_local": "2026-09-23 22:17:06 +0800",
|
||||||
|
"python": "3.14.7 (main, Aug 10 2026, 00:00:00) [GCC 16.1.1 20260515 (Red Hat 16.1.1-2)]",
|
||||||
|
"platform": "Linux-6.6.87.2-microsoft-standard-WSL2-x86_64-with-glibc2.43",
|
||||||
|
"uv_version": "uv 0.11.28 (x86_64-unknown-linux-gnu)",
|
||||||
|
"packages": {
|
||||||
|
"torch": "2.14.0+cu130",
|
||||||
|
"transformers": "5.17.0",
|
||||||
|
"numpy": "2.5.3",
|
||||||
|
"scipy": "1.18.1",
|
||||||
|
"scikit-learn": "1.9.1",
|
||||||
|
"mediapipe": "1.0.1",
|
||||||
|
"opencv-python-headless": "5.0.0.93",
|
||||||
|
"openpyxl": "3.1.5",
|
||||||
|
"matplotlib": "3.11.2"
|
||||||
|
},
|
||||||
|
"device": "cuda",
|
||||||
|
"cuda_available": true,
|
||||||
|
"gpu_name": "NVIDIA GeForce RTX 5070 Ti",
|
||||||
|
"models": {
|
||||||
|
"text_id": "google-bert/bert-base-uncased",
|
||||||
|
"text_revision": "86b5e0934494bd15c9632b12f734a8a67f723594",
|
||||||
|
"speech_id": "facebook/wav2vec2-base-960h",
|
||||||
|
"speech_revision": "22aad52d435eb6dbaf354bdad9b0da84ce7d6156",
|
||||||
|
"speech_stride_samples": 320,
|
||||||
|
"speech_receptive_field_samples": 400
|
||||||
|
},
|
||||||
|
"face_landmarker_asset_sha256": "64184e229b263107bc2b804c6625db1341ff2bb731874b0bcc2fe6544e0bc9ff",
|
||||||
|
"inputs": {
|
||||||
|
"label_file": "/home/gloamxun/modeling_zhaocui/E题数据/附件1-数据集原始多模态样本/MOSEI数据集部分原始视频-100条/label-100.xlsx",
|
||||||
|
"label_file_sha256": "827334f782b1f242c84ad944a657bc7f7c63643f71ae9c977811a3eecdd3ab33",
|
||||||
|
"sample_count": 100,
|
||||||
|
"video_sha256_by_sample": {
|
||||||
|
"-3g5yACwYnA/13": "aeb47627f59dce3d0a85a44ef35e4a3bc18211498e98c8c11cf60269646df24f",
|
||||||
|
"-3g5yACwYnA/3": "eff8cfefba2425155f2b72656829b34b23be8bfe1dcbece384154f56818aa26d",
|
||||||
|
"-3g5yACwYnA/2": "0619c01f137d017c15c058176f18a575919189c7e81ebd2bfb993afdb86f3de7",
|
||||||
|
"-3g5yACwYnA/9": "b962573b12ed1f06d5533f6427c80dce2bae3a99dd332f6d3fcb32c579f3f016",
|
||||||
|
"-3nNcZdcdvU/5": "7ea53ad502b77be7d2ee4494dc23a9478c897792203cf887d546e83d584c1728",
|
||||||
|
"-HwX2H8Z4hY/2": "920a1052c09ebd1eedba6dd1ccfecd80f56f6f0fdcdd1f6dd32b0d90bcb89358",
|
||||||
|
"-HwX2H8Z4hY/5": "8d61b7f5840a8bd403d9bc6c3cd49029ffc4e304cf9c7834735c1a6bb5fd369d",
|
||||||
|
"-HwX2H8Z4hY/6": "696576dcbd0dc42cb678fa40d8a0f0f419a3072bfa7662d653756ab29bba4c46",
|
||||||
|
"-HwX2H8Z4hY/9": "90c18e627f27a12c14e21ba24c79edf3008762b4ac26874b490ad3f888381702",
|
||||||
|
"-NFrJFQijFE/1": "d8fcf16c6eb51c56947ec3d0549c6f3bc7a0cf1fc5f5d89e8b38d4249a08db4f",
|
||||||
|
"-NFrJFQijFE/2": "a54a5c144a64f5abd26b19aafa6cea8142ee08e9385fa508fa70c934f6b616c1",
|
||||||
|
"-THoVjtIkeU/12": "e2f148f4f2e74724dc13b3ce870a0ce6a06740cfc224f5a9b3d0a870b196e7d4",
|
||||||
|
"-THoVjtIkeU/2": "35a6c2464ffd68dd96411edb5dc2763f0efb59f264c1d8b737e616fa6b0d1d5b",
|
||||||
|
"-THoVjtIkeU/6": "efaa55ee6394032c867046690a92c29edfd825005f0a090cca5ccb763b227f64",
|
||||||
|
"-UuX1xuaiiE/1": "d5bbda38fcec15817d2b87bab5dcc559d6d425f7d28a82e6c32d13f14b48650c",
|
||||||
|
"-UuX1xuaiiE/0": "f578b305d53bc152702f5e771de0fcd12d1e9aaefc5cefce3d2082a417aab7c8",
|
||||||
|
"-UuX1xuaiiE/3": "eddb408f25bdc13e6de2fb74caeed709cb7a8f00640e7897135330437f66a2aa",
|
||||||
|
"-UuX1xuaiiE/6": "f0131ac00e44410fbd32c547a0d421c2791172434fc0203bb969abe14a530532",
|
||||||
|
"-a55Q6RWvTA/3": "15d029fc15f50b268b98f1e8abc65e45582e638c13a018d7aa74a18373275946",
|
||||||
|
"-aNfi7CP8vM/7": "0b6389f45bb966113c65a11998c6e7facdd24c8c42acf2e8cf32e0f2139402e1",
|
||||||
|
"-aqamKhZ1Ec/0": "6f6e674bbba4353399e9e7825217f37e6d8ca1adf9675f33cf37106f31c55548",
|
||||||
|
"-dxfTGcXJoc/1": "b5ffdc98a4a98b8f0aae55dee5fed66a43bcb129f47a250a004c1cc74e572595",
|
||||||
|
"-dxfTGcXJoc/0": "04bd0be907fb46c613f926fe1b06bc2c3b30881c293156510baa34a4105ade55",
|
||||||
|
"-dxfTGcXJoc/2": "46a5e523b5dd00364c99a3bcbc6fae4c80e1382a84b5b5e681c46eb074ed10f7",
|
||||||
|
"-dxfTGcXJoc/6": "55c1ff86729ab21da743a5107ed07b220e41b4193cd1c0ba9f009e65c3ab0513",
|
||||||
|
"-egA8-b7-3M/26": "11f78aed680ea9c5862a15e04ff6c1f8776c46d29c98b68d28773ff39d959685",
|
||||||
|
"-egA8-b7-3M/17": "f967e81e0edd51c3700671721afb3fbecbbc32652f27d2db40ea444e8d80c68f",
|
||||||
|
"-egA8-b7-3M/18": "fd4bc95a6adfb16a9588b7ed65cc2812186ce3d607804dc76f4486312951ef90",
|
||||||
|
"-egA8-b7-3M/16": "4c015f85e3bd901ecedde8c0c2ca4a756d9a77dfe4177625c481581f2037ff9c",
|
||||||
|
"-egA8-b7-3M/13": "cbd627c225ebca37521ae238cb9ec254cad236822f4f37ebfbb039e6c56aea4c",
|
||||||
|
"-egA8-b7-3M/1": "589a7989b1b4888568c31614a2d6beefba88fb87289c85b94df2d8051500403b",
|
||||||
|
"-egA8-b7-3M/6": "2e88a00d5e55863aa956b4b0949fc0e4f2975ccfd83b3194957f6fca86ecfce0",
|
||||||
|
"-egA8-b7-3M/9": "bcd643f328616312eae9b28e6c9aabbe8acaec3e4c4be58daeb45e3eee4d973c",
|
||||||
|
"-egA8-b7-3M/20": "719ef133920e74b3ea33925d2bca0270b062199116867568f88ac83cadcd74e8",
|
||||||
|
"-iRBcNs9oI8/3": "39c547acd1a8da2ebc6c6ab008191a5676420398a974cd9611cad087f0ccec50",
|
||||||
|
"-iRBcNs9oI8/7": "cc7b1c7a06ec41d72007b5c42fb8036f30b66d5cf76e3f093160c22fc5e38081",
|
||||||
|
"-iRBcNs9oI8/6": "030b0b8d8d4ae47799432ab71b66fcc881f077c5b0feef892c1f146dc0a592e3",
|
||||||
|
"-iRBcNs9oI8/9": "352bdcc73d3622d6abac1b09f2039b211d15b9bf7acee66f0b07783c9b7a4c74",
|
||||||
|
"-iRBcNs9oI8/8": "72a49e3da2862addcde033bd2fbae957d086533664ca4dcbccd3a851bcd50412",
|
||||||
|
"-lzEya4AM_4/5": "3b74de8d0e66fd754594f05985e5af3480adcc03555f79225d3b727cb4bed043",
|
||||||
|
"-lzEya4AM_4/6": "ff5da5afb551001d58100ed2471caea51c741046d5396957efbf40082307ca18",
|
||||||
|
"-mJ2ud6oKI8/1": "c8e3acb4ca08679796c4c8cdfcd8c2b3cffc0465015efaffa1eccdb55ff4f45b",
|
||||||
|
"-mJ2ud6oKI8/2": "c41790f8b75f886bd1f5d6d72cbab1985306d0b715e5ae47203f3d0c8a3e7b28",
|
||||||
|
"-mJ2ud6oKI8/6": "134a3a4ebbc423760133b0da0e90cfe84d75ea85df3968000afd84153c5d828d",
|
||||||
|
"-mJ2ud6oKI8/9": "df3442f3ed8f894b507923e87e426ae9636790a4a502d00c350a215270cc3e47",
|
||||||
|
"-mJ2ud6oKI8/8": "6af6614c9838417a48c4817f4dd5bf1e6669eb5f78f983e9a68adc113fcdc7cc",
|
||||||
|
"-mqbVkbCndg/0": "cdeba953bba30907814acf72f3e35582dec263eaa25bdd922dad7fafe02e4dda",
|
||||||
|
"-t217m2on-s/2": "6713a983112173204502dc5adbb886bc6c38554b4627698611931420c47663a8",
|
||||||
|
"-t217m2on-s/7": "41c76b8a733ecb780d56348f55d00f0ba0c07f44f46e93044e62d876710180b7",
|
||||||
|
"-tANM6ETl_M/3": "6d3af064c060dad3816a9e1dfa00101faebd8b7da7d8ceea85b0bc0fca70abd6",
|
||||||
|
"-tPCytz4rww/11": "6dc09d02467baaef67594ff32cffb01f8ed1565cda4c5abd161b8261a51e2e5d",
|
||||||
|
"-tPCytz4rww/10": "758183192cbd5f0f20823d26ece3ddee40dec0aadec35247b2495510a53f52f8",
|
||||||
|
"-tPCytz4rww/12": "e3c7f9cd67ad2d20fef997696dc278361f97f027db58762e22104e5720d509d5",
|
||||||
|
"-tPCytz4rww/16": "8b3f82f9628792eebaf8c227becc2ebeba59bb8abb668bcca86e07fb1a992a93",
|
||||||
|
"-tPCytz4rww/18": "eece70451e9c37a4b5f375bf646f66e09113fe3467f35bf1720cc14d986c1776",
|
||||||
|
"-vxjVxOeScU/4": "00282df9a314394f19d616d561bd1916d1d30dca4c814dfb43630ca164f4c719",
|
||||||
|
"-wMB_hJL-3o/7": "04a73d73fd150b07edab8e652fd14c9cb4a0f6e017ea97b17e8146e377b9fae7",
|
||||||
|
"-wny0OAz3g8/1": "c5535859f129ce04da5f5b5104d02d4168057b9d947a5349d423278ad4f3d8e8",
|
||||||
|
"-wny0OAz3g8/0": "1cc21b4432e2f4b592284ee40558a18acc7d208019117aebdb2f6a1b67ea198d",
|
||||||
|
"-wny0OAz3g8/3": "ea917c506bc9fa39460f2b722f7c416f646312bb17cd74baa618905b54837bce",
|
||||||
|
"-wny0OAz3g8/2": "1825fb37db2e914fb616cd80aebd613afd873eecb7edd4d0b53158bfda0e522a",
|
||||||
|
"-wny0OAz3g8/5": "6c6a8a4cba654b25190ab1793bfae81c3d2159ecd6cbc02d3fcbd598543acce0",
|
||||||
|
"-wny0OAz3g8/7": "cbf9cac1048ce191aef781d0c4563eb97f9f3ddd52eb1f9551fa78d31d70e7cd",
|
||||||
|
"-wny0OAz3g8/9": "b76f70fe783ee76bf1b637137226c98e67d13a6b15ec725894175b36cd76a136",
|
||||||
|
"-571d8cVauQ/0": "defc62e1530d9b0cd27f2acfe042b71dd84d7e8c85f1df261f04912091976088",
|
||||||
|
"-571d8cVauQ/5": "e304c03fac1f477823fe0926c4fb421b9c4e4091f85718b729afe856c24a7d3e",
|
||||||
|
"-I_e4mIh0yE/1": "8453a966d4bd1f0179eb86a35cd34c0864a3a3bc2fa86f37783060cc0d14f5df",
|
||||||
|
"-I_e4mIh0yE/3": "a015f02ae51d74c42ad27e2a0831263a303f0658ac8d984819e27ad61342b23e",
|
||||||
|
"-UacrmKiTn4/10": "00312e0602dbeac970de5d7dd0e8a970e514fe45d7fe6625b4f454b299e36360",
|
||||||
|
"-UacrmKiTn4/4": "a7f2ab0d6ee2e2a2d5cd4239f03e6cac67c618ddb744891323df3ff87e29af84",
|
||||||
|
"-hnBHBN8p5A/7": "0ec6fb76226315d3fafbb483e3500e60cc8a6729617b10e4b4b76163595080fb",
|
||||||
|
"-hnBHBN8p5A/6": "997f6808ac3973ddd71800e91c8e10755c10b5037186d84348a79e00d672fdc7",
|
||||||
|
"-qDkUB0GgYY/6": "ab19bf4a761d3919b7113b1104ccaadbbf0c4e7b45694904d094280a2c329c8b",
|
||||||
|
"-uywlfIYOS8/4": "fde46f5222cf5cfec7af9c93214c972ce2d686b0bf9c90a291021601968986f9",
|
||||||
|
"-6rXp3zJ3kc/8": "59b77a0f873b7fa95e7327b93d64bb5a4bdae84039321988ef70bba8e63ee1ae",
|
||||||
|
"-9y-fZ3swSY/0": "29a1fd181293cdbe24e50edcad51e3cc0f27a5824d840bcc39d453b272fb6e3c",
|
||||||
|
"-9y-fZ3swSY/4": "fe04a26710e85c58bdaad5a48618c7e0f742d82b0d920eb4a046c8f14a6921dc",
|
||||||
|
"-9y-fZ3swSY/8": "6140d031e614b62a5d8df73346fef2417c28cee3501310aa95aa40d358b4b256",
|
||||||
|
"-AUZQgSxyPQ/2": "e09034b30bb4f8fb9f9c2568afc20f2a764e327fe6ef4e13f21804eaac51bb7b",
|
||||||
|
"-HeZS2-Prhc/2": "edb0fa22fdfbd990c8174764af78670bd2ecad799240d35d96a1124a3cd9dc0c",
|
||||||
|
"-MeTTeMJBNc/0": "54e2097f7bdb455d243e72254e5af7ed16160739f2f3cfd22f57d2729ec1acbe",
|
||||||
|
"-MeTTeMJBNc/13": "80c5d996e5d7b61c58e1fe32d706efcaf80c8c8cddb5c37a5c66d527786c2a44",
|
||||||
|
"-MeTTeMJBNc/7": "40716dc7402fd6398038eded58ce939ec76be6afbeea346b968bf4c427088f3a",
|
||||||
|
"-RfYyzHpjk4/11": "96e481cd732001239639c8cb4e7f94e570925e5a15940920ffab0cec1f13f904",
|
||||||
|
"-RfYyzHpjk4/8": "ed8c672c532b8c7d1c830c82d6e2c14fcb3bbd3dc1a86670a03a6c8e83c02f3c",
|
||||||
|
"-RfYyzHpjk4/2": "5e428ee22de58c9d6561b9b4353bef9dfb811a5841d9ab07b5bc453d697ad17f",
|
||||||
|
"-UUCSKoHeMA/0": "3f9792888ec8e969f61cfb6cb4de7bdf7ef8944afe0a9d2a9d13586e1f13b898",
|
||||||
|
"-ri04Z7vwnc/0": "93021c70c8fad20ad3fded80e7fc790a2f9c96835de58a3b694c6906de42684b",
|
||||||
|
"-ri04Z7vwnc/2": "f5993ce3a6ee56298462c08692a96f78c7d914d7264257ed401de7d6e17fd148",
|
||||||
|
"-ri04Z7vwnc/5": "2a5e4319bf592a18c8eb5117f7c7ec30c6c79528f9eac9fe4fb85bf70f277736",
|
||||||
|
"-s9qJ7ATP7w/1": "a33e8f82a185bcb7648b523927308cdc2142e3076479ba2b45cd0e828636fc09",
|
||||||
|
"-s9qJ7ATP7w/0": "f2324dbc6debe0d1e4254e6f943523373e3cbb30d8d0555413ce3a7b370a4af5",
|
||||||
|
"-s9qJ7ATP7w/5": "bd35eb8de64ce1f8b27921294598f7cf5bb971255738f51d06458c7096c58eae",
|
||||||
|
"-s9qJ7ATP7w/4": "76e0672ffce5857a789b41fb10606402f7da58c33579b1032a8282f86fb35bc9",
|
||||||
|
"-s9qJ7ATP7w/7": "5301f4403cf7cd299cc525d2f443392ac70fd70c3c9cd4a09406f2ba4df6734d",
|
||||||
|
"-s9qJ7ATP7w/6": "40b684fd1b4559f0f82e72f9b47f4ce7c5e15059fe8d0d8ab112637cd3386451",
|
||||||
|
"-s9qJ7ATP7w/8": "1de4d91bd9d9fb508a3779658747da88460db0bdb1beefe7ea5b2c77edfd2239",
|
||||||
|
"-yRb-Jum7EQ/1": "2f16da1baa8660e5bfc608f5237cdd59725d16e57a19bc02c4556cd2c7129a46",
|
||||||
|
"-yRb-Jum7EQ/5": "6d6145293cf84d1fc8324dfa8d39465e2d7fc07affb67ff68deeabf6e3c410d9",
|
||||||
|
"-yRb-Jum7EQ/6": "b51eb9ad06de804934cb6a315de591ba9ff7f8aa71bffff7f8001ab29945d989"
|
||||||
|
},
|
||||||
|
"label_polarity_consistency_count": 100
|
||||||
|
},
|
||||||
|
"experiment": {
|
||||||
|
"methods": {
|
||||||
|
"B0": "quality-weighted hard CTC word projection + physical source-time pooling + train-fold standard scaling",
|
||||||
|
"B1": "B0 + train-fold median/MAD scaling + SO(3) rotation mean + normalized gaze vector mean",
|
||||||
|
"B2": "B1 + fixed-transcript CTC forward-backward word occupancy projection for text",
|
||||||
|
"B3": "B1 + second-order time-augmented log signatures of selected audio/face channels",
|
||||||
|
"B4": "B1 + frozen Wav2Vec2 last-four-layer speech representation"
|
||||||
|
},
|
||||||
|
"text_dimension": 768,
|
||||||
|
"audio_dimension": 74,
|
||||||
|
"vision_dimension": 35,
|
||||||
|
"audio_features": [
|
||||||
|
"logmel_00",
|
||||||
|
"logmel_01",
|
||||||
|
"logmel_02",
|
||||||
|
"logmel_03",
|
||||||
|
"logmel_04",
|
||||||
|
"logmel_05",
|
||||||
|
"logmel_06",
|
||||||
|
"logmel_07",
|
||||||
|
"logmel_08",
|
||||||
|
"logmel_09",
|
||||||
|
"logmel_10",
|
||||||
|
"logmel_11",
|
||||||
|
"logmel_12",
|
||||||
|
"logmel_13",
|
||||||
|
"logmel_14",
|
||||||
|
"logmel_15",
|
||||||
|
"logmel_16",
|
||||||
|
"logmel_17",
|
||||||
|
"logmel_18",
|
||||||
|
"logmel_19",
|
||||||
|
"logmel_20",
|
||||||
|
"logmel_21",
|
||||||
|
"logmel_22",
|
||||||
|
"logmel_23",
|
||||||
|
"logmel_24",
|
||||||
|
"logmel_25",
|
||||||
|
"logmel_26",
|
||||||
|
"logmel_27",
|
||||||
|
"logmel_28",
|
||||||
|
"logmel_29",
|
||||||
|
"logmel_30",
|
||||||
|
"logmel_31",
|
||||||
|
"logmel_32",
|
||||||
|
"logmel_33",
|
||||||
|
"logmel_34",
|
||||||
|
"logmel_35",
|
||||||
|
"logmel_36",
|
||||||
|
"logmel_37",
|
||||||
|
"logmel_38",
|
||||||
|
"logmel_39",
|
||||||
|
"mfcc_00",
|
||||||
|
"mfcc_01",
|
||||||
|
"mfcc_02",
|
||||||
|
"mfcc_03",
|
||||||
|
"mfcc_04",
|
||||||
|
"mfcc_05",
|
||||||
|
"mfcc_06",
|
||||||
|
"mfcc_07",
|
||||||
|
"mfcc_08",
|
||||||
|
"mfcc_09",
|
||||||
|
"mfcc_10",
|
||||||
|
"mfcc_11",
|
||||||
|
"mfcc_12",
|
||||||
|
"delta_mfcc_00",
|
||||||
|
"delta_mfcc_01",
|
||||||
|
"delta_mfcc_02",
|
||||||
|
"delta_mfcc_03",
|
||||||
|
"delta_mfcc_04",
|
||||||
|
"delta_mfcc_05",
|
||||||
|
"delta_mfcc_06",
|
||||||
|
"delta_mfcc_07",
|
||||||
|
"delta_mfcc_08",
|
||||||
|
"delta_mfcc_09",
|
||||||
|
"delta_mfcc_10",
|
||||||
|
"delta_mfcc_11",
|
||||||
|
"delta_mfcc_12",
|
||||||
|
"log_energy",
|
||||||
|
"log_f0_hz",
|
||||||
|
"voicing_strength",
|
||||||
|
"spectral_centroid_hz",
|
||||||
|
"spectral_bandwidth_hz",
|
||||||
|
"spectral_flux",
|
||||||
|
"zero_crossing_rate",
|
||||||
|
"hnr_db"
|
||||||
|
],
|
||||||
|
"vision_features": [
|
||||||
|
"mp_blendshape_browDownLeft",
|
||||||
|
"mp_blendshape_browDownRight",
|
||||||
|
"mp_blendshape_browInnerUp",
|
||||||
|
"mp_blendshape_browOuterUpLeft",
|
||||||
|
"mp_blendshape_browOuterUpRight",
|
||||||
|
"mp_blendshape_eyeBlinkLeft",
|
||||||
|
"mp_blendshape_eyeBlinkRight",
|
||||||
|
"mp_blendshape_eyeSquintLeft",
|
||||||
|
"mp_blendshape_eyeSquintRight",
|
||||||
|
"mp_blendshape_eyeWideLeft",
|
||||||
|
"mp_blendshape_eyeWideRight",
|
||||||
|
"mp_blendshape_jawOpen",
|
||||||
|
"mp_blendshape_mouthFrownLeft",
|
||||||
|
"mp_blendshape_mouthFrownRight",
|
||||||
|
"mp_blendshape_mouthPucker",
|
||||||
|
"mp_blendshape_mouthSmileLeft",
|
||||||
|
"mp_blendshape_mouthSmileRight",
|
||||||
|
"pose_rotvec_x",
|
||||||
|
"pose_rotvec_y",
|
||||||
|
"pose_rotvec_z",
|
||||||
|
"pose_translation_x",
|
||||||
|
"pose_translation_y",
|
||||||
|
"pose_translation_z",
|
||||||
|
"gaze_left_x",
|
||||||
|
"gaze_left_y",
|
||||||
|
"gaze_left_z",
|
||||||
|
"gaze_right_x",
|
||||||
|
"gaze_right_y",
|
||||||
|
"gaze_right_z",
|
||||||
|
"eye_aperture_left",
|
||||||
|
"eye_aperture_right",
|
||||||
|
"mouth_aperture",
|
||||||
|
"mouth_width_face_ratio",
|
||||||
|
"brow_eye_distance_left",
|
||||||
|
"brow_eye_distance_right"
|
||||||
|
],
|
||||||
|
"vision_feature_note": "17 MediaPipe blendshape proxies and approximate gaze/landmark geometry; not OpenFace AU labels",
|
||||||
|
"grid_step_s": 0.1,
|
||||||
|
"audio_window_samples": 400,
|
||||||
|
"audio_step_samples": 160,
|
||||||
|
"audio_fft": 512,
|
||||||
|
"audio_mel_bands": 40,
|
||||||
|
"vision_sample_rate_hz": 5.0,
|
||||||
|
"grouping": "GroupKFold by video_id",
|
||||||
|
"fold_count": 5,
|
||||||
|
"probe": {
|
||||||
|
"classifier": "LogisticRegression",
|
||||||
|
"C": 0.05,
|
||||||
|
"regressor": "Ridge",
|
||||||
|
"alpha": 25.0,
|
||||||
|
"intensity_clip": [
|
||||||
|
-3.0,
|
||||||
|
3.0
|
||||||
|
]
|
||||||
|
},
|
||||||
|
"bootstrap": {
|
||||||
|
"unit": "video_id",
|
||||||
|
"repeats": 2000,
|
||||||
|
"seed": 20260923
|
||||||
|
},
|
||||||
|
"human_boundary_reference_count": 0,
|
||||||
|
"boundary_interval_level": 0.9,
|
||||||
|
"hard_text_quality_weight": "per-sample median-normalized CTC path score clipped to [0.25, 4.0]; uncalibrated",
|
||||||
|
"cache_schema": "q1-b0b4-v3"
|
||||||
|
},
|
||||||
|
"output_dir": "/home/gloamxun/modeling_zhaocui/math/results/model_comparison",
|
||||||
|
"typical_sample_id": "-9y-fZ3swSY/0",
|
||||||
|
"elapsed_seconds": 124.10033178329468
|
||||||
|
}
|
||||||
@@ -0,0 +1,101 @@
|
|||||||
|
sample_id,video_id,clip_id,source_video_sha256,duration_s,word_count,hard_aligned_word_count,unlocated_word_count,posterior_usable_word_count,mean_start_interval_width90_s,mean_end_interval_width90_s,text_hard_grid_coverage,text_posterior_grid_occupancy,audio_grid_coverage,vision_grid_coverage,audio_native_rows,vision_native_rows,status
|
||||||
|
-3g5yACwYnA/13,-3g5yACwYnA,13,aeb47627f59dce3d0a85a44ef35e4a3bc18211498e98c8c11cf60269646df24f,5.5139970779418945,15,15,0,15,0.037333333333333406,0.020000000000000004,0.7015532851219177,0.24074573814868927,0.9882608652114868,1.0,540,43,ok
|
||||||
|
-3g5yACwYnA/3,-3g5yACwYnA,3,eff8cfefba2425155f2b72656829b34b23be8bfe1dcbece384154f56818aa26d,14.388997077941895,29,29,0,29,0.2793103448275863,0.25931034482758614,0.6343300342559814,0.19729532301425934,0.9912122488021851,1.0,1433,103,ok
|
||||||
|
-3g5yACwYnA/2,-3g5yACwYnA,2,0619c01f137d017c15c058176f18a575919189c7e81ebd2bfb993afdb86f3de7,9.394009590148926,14,14,0,14,0.2828571428571428,0.19142857142857145,0.7563282251358032,0.15269099175930023,0.9936367273330688,1.0,924,78,ok
|
||||||
|
-3g5yACwYnA/9,-3g5yACwYnA,9,b962573b12ed1f06d5533f6427c80dce2bae3a99dd332f6d3fcb32c579f3f016,8.816991806030273,21,21,0,21,0.22095238095238093,0.1942857142857142,0.6637077927589417,0.26207172870635986,0.9912108778953552,1.0,870,75,ok
|
||||||
|
-3nNcZdcdvU/5,-3nNcZdcdvU,5,7ea53ad502b77be7d2ee4494dc23a9478c897792203cf887d546e83d584c1728,7.867969036102295,18,18,0,18,0.03555555555555544,0.012222222222222258,0.7357116341590881,0.22051197290420532,0.9971166849136353,0.9904456734657288,779,67,ok
|
||||||
|
-HwX2H8Z4hY/2,-HwX2H8Z4hY,2,920a1052c09ebd1eedba6dd1ccfecd80f56f6f0fdcdd1f6dd32b0d90bcb89358,3.9820311069488525,8,7,1,7,0.12285714285714287,0.12285714285714275,0.6164926290512085,0.18007618188858032,0.9881895780563354,0.9814813733100891,392,27,ok
|
||||||
|
-HwX2H8Z4hY/5,-HwX2H8Z4hY,5,8d61b7f5840a8bd403d9bc6c3cd49029ffc4e304cf9c7834735c1a6bb5fd369d,5.6529951095581055,10,9,1,9,0.037777777777777764,0.040000000000000036,0.7021505832672119,0.22856950759887695,0.9900854229927063,0.666666567325592,555,44,ok
|
||||||
|
-HwX2H8Z4hY/6,-HwX2H8Z4hY,6,696576dcbd0dc42cb678fa40d8a0f0f419a3072bfa7662d653756ab29bba4c46,3.3580079078674316,7,7,0,7,0.048571428571428564,0.03428571428571422,0.5562952756881714,0.1454559713602066,0.9901678562164307,0.841666579246521,326,21,ok
|
||||||
|
-HwX2H8Z4hY/9,-HwX2H8Z4hY,9,90c18e627f27a12c14e21ba24c79edf3008762b4ac26874b490ad3f888381702,7.733983993530273,22,22,0,22,0.16363636363636375,0.15181818181818188,0.6705501675605774,0.25460928678512573,0.9946621060371399,0.0,764,65,ok
|
||||||
|
-NFrJFQijFE/1,-NFrJFQijFE,1,d8fcf16c6eb51c56947ec3d0549c6f3bc7a0cf1fc5f5d89e8b38d4249a08db4f,5.745999813079834,16,16,0,16,0.46625000000000005,0.47125000000000006,0.7835317254066467,0.27368244528770447,0.9922494292259216,0.0,566,44,ok
|
||||||
|
-NFrJFQijFE/2,-NFrJFQijFE,2,a54a5c144a64f5abd26b19aafa6cea8142ee08e9385fa508fa70c934f6b616c1,6.855999946594238,18,18,0,18,0.17888888888888899,0.1544444444444445,0.7431879639625549,0.24775999784469604,0.9926568269729614,0.0,670,55,ok
|
||||||
|
-THoVjtIkeU/12,-THoVjtIkeU,12,e2f148f4f2e74724dc13b3ce870a0ce6a06740cfc224f5a9b3d0a870b196e7d4,14.896029472351074,39,39,0,39,0.05999999999999991,0.05282051282051286,0.6352806687355042,0.23379945755004883,0.9911767244338989,1.0,1481,106,ok
|
||||||
|
-THoVjtIkeU/2,-THoVjtIkeU,2,35a6c2464ffd68dd96411edb5dc2763f0efb59f264c1d8b737e616fa6b0d1d5b,4.2919921875,10,10,0,10,0.016000000000000004,0.007999999999999985,0.7164996266365051,0.2465074211359024,0.9895758628845215,1.0,424,31,ok
|
||||||
|
-THoVjtIkeU/6,-THoVjtIkeU,6,efaa55ee6394032c867046690a92c29edfd825005f0a090cca5ccb763b227f64,8.097004890441895,30,30,0,30,0.053333333333333295,0.05733333333333333,0.6787540316581726,0.30513063073158264,0.9927164316177368,1.0,799,68,ok
|
||||||
|
-UuX1xuaiiE/1,-UuX1xuaiiE,1,d5bbda38fcec15817d2b87bab5dcc559d6d425f7d28a82e6c32d13f14b48650c,10.350000381469727,27,27,0,27,0.09555555555555556,0.060000000000000046,0.6550575494766235,0.22912898659706116,0.9925051927566528,0.8333339691162109,1027,83,ok
|
||||||
|
-UuX1xuaiiE/0,-UuX1xuaiiE,0,f578b305d53bc152702f5e771de0fcd12d1e9aaefc5cefce3d2082a417aab7c8,3.6333329677581787,9,9,0,9,0.03777777777777775,0.04666666666666665,0.6484582424163818,0.2555619180202484,0.994590699672699,0.9444444179534912,358,25,ok
|
||||||
|
-UuX1xuaiiE/3,-UuX1xuaiiE,3,eddb408f25bdc13e6de2fb74caeed709cb7a8f00640e7897135330437f66a2aa,4.1300129890441895,9,9,0,9,0.07777777777777779,0.0600000000000001,0.7240580320358276,0.16585363447666168,0.992123007774353,0.976190447807312,404,29,ok
|
||||||
|
-UuX1xuaiiE/6,-UuX1xuaiiE,6,f0131ac00e44410fbd32c547a0d421c2791172434fc0203bb969abe14a530532,8.113997459411621,22,22,0,22,0.06363636363636366,0.0736363636363637,0.7064024806022644,0.25750458240509033,0.984533965587616,0.9776423573493958,795,68,ok
|
||||||
|
-a55Q6RWvTA/3,-a55Q6RWvTA,3,15d029fc15f50b268b98f1e8abc65e45582e638c13a018d7aa74a18373275946,22.15397071838379,65,65,0,65,0.2215384615384614,0.18892307692307703,0.6580277681350708,0.23805321753025055,0.9912921786308289,1.0,2209,142,ok
|
||||||
|
-aNfi7CP8vM/7,-aNfi7CP8vM,7,0b6389f45bb966113c65a11998c6e7facdd24c8c42acf2e8cf32e0f2139402e1,8.694987297058105,18,18,0,18,0.2511111111111111,0.24000000000000002,0.7001731395721436,0.23255421221256256,0.9945195913314819,1.0,854,74,ok
|
||||||
|
-aqamKhZ1Ec/0,-aqamKhZ1Ec,0,6f6e674bbba4353399e9e7825217f37e6d8ca1adf9675f33cf37106f31c55548,10.966667175292969,14,13,1,13,0.22461538461538455,0.08615384615384586,0.71061772108078,0.1247716099023819,0.9910455346107483,1.0,1090,86,ok
|
||||||
|
-dxfTGcXJoc/1,-dxfTGcXJoc,1,b5ffdc98a4a98b8f0aae55dee5fed66a43bcb129f47a250a004c1cc74e572595,16.16100311279297,36,36,0,36,0.07777777777777778,0.033333333333333375,0.6812966465950012,0.2449914664030075,0.9900091290473938,1.0,1602,112,ok
|
||||||
|
-dxfTGcXJoc/0,-dxfTGcXJoc,0,04bd0be907fb46c613f926fe1b06bc2c3b30881c293156510baa34a4105ade55,20.5,41,38,3,38,0.09315789473684201,0.06789473684210537,0.7333337664604187,0.209048330783844,0.9891952872276306,0.99895840883255,2042,134,ok
|
||||||
|
-dxfTGcXJoc/2,-dxfTGcXJoc,2,46a5e523b5dd00364c99a3bcbc6fae4c80e1382a84b5b5e681c46eb074ed10f7,16.697982788085938,34,32,2,32,0.048124999999999946,0.03624999999999996,0.6593995094299316,0.2266717106103897,0.9901677966117859,1.0,1664,115,ok
|
||||||
|
-dxfTGcXJoc/6,-dxfTGcXJoc,6,55c1ff86729ab21da743a5107ed07b220e41b4193cd1c0ba9f009e65c3ab0513,12.735026359558105,27,27,0,27,0.05851851851851849,0.05407407407407409,0.715381383895874,0.25197115540504456,0.9889141917228699,1.0,1263,95,ok
|
||||||
|
-egA8-b7-3M/26,-egA8-b7-3M,26,11f78aed680ea9c5862a15e04ff6c1f8776c46d29c98b68d28773ff39d959685,6.030990123748779,15,15,0,15,0.08133333333333341,0.02533333333333337,0.6877890825271606,0.22332797944545746,0.9945786595344543,1.0,593,47,ok
|
||||||
|
-egA8-b7-3M/17,-egA8-b7-3M,17,f967e81e0edd51c3700671721afb3fbecbbc32652f27d2db40ea444e8d80c68f,7.271028995513916,16,16,0,16,0.017499999999999995,0.019999999999999962,0.7292676568031311,0.24067606031894684,0.9916234016418457,1.0,710,60,ok
|
||||||
|
-egA8-b7-3M/18,-egA8-b7-3M,18,fd4bc95a6adfb16a9588b7ed65cc2812186ce3d607804dc76f4486312951ef90,8.386002540588379,22,22,0,22,0.052727272727272734,0.021818181818181796,0.6243576407432556,0.26077014207839966,0.9932008385658264,1.0,827,71,ok
|
||||||
|
-egA8-b7-3M/16,-egA8-b7-3M,16,4c015f85e3bd901ecedde8c0c2ca4a756d9a77dfe4177625c481581f2037ff9c,5.538021087646484,11,11,0,11,0.010909090909090908,0.0181818181818182,0.7375588417053223,0.23636400699615479,0.9940068125724792,1.0,549,43,ok
|
||||||
|
-egA8-b7-3M/13,-egA8-b7-3M,13,cbd627c225ebca37521ae238cb9ec254cad236822f4f37ebfbb039e6c56aea4c,4.18398380279541,11,11,0,11,0.23818181818181816,0.23090909090909092,0.6372604370117188,0.2536529004573822,0.9924017786979675,1.0,403,29,ok
|
||||||
|
-egA8-b7-3M/1,-egA8-b7-3M,1,589a7989b1b4888568c31614a2d6beefba88fb87289c85b94df2d8051500403b,10.266016006469727,20,19,1,19,1.0621052631578947,1.0526315789473684,0.6490206122398376,0.19601602852344513,0.9937204718589783,1.0,1016,83,ok
|
||||||
|
-egA8-b7-3M/6,-egA8-b7-3M,6,2e88a00d5e55863aa956b4b0949fc0e4f2975ccfd83b3194957f6fca86ecfce0,6.264974117279053,12,12,0,12,0.08833333333333332,0.03666666666666665,0.7164138555526733,0.2193591147661209,0.9942464828491211,1.0,619,50,ok
|
||||||
|
-egA8-b7-3M/9,-egA8-b7-3M,9,bcd643f328616312eae9b28e6c9aabbe8acaec3e4c4be58daeb45e3eee4d973c,5.0899739265441895,10,10,0,10,0.007999999999999962,0.011999999999999966,0.597113311290741,0.2590937316417694,0.9939422011375427,1.0,494,38,ok
|
||||||
|
-egA8-b7-3M/20,-egA8-b7-3M,20,719ef133920e74b3ea33925d2bca0270b062199116867568f88ac83cadcd74e8,6.644987106323242,20,20,0,20,0.077,0.021000000000000008,0.6985589861869812,0.2369300276041031,0.992087185382843,1.0,648,54,ok
|
||||||
|
-iRBcNs9oI8/3,-iRBcNs9oI8,3,39c547acd1a8da2ebc6c6ab008191a5676420398a974cd9611cad087f0ccec50,5.620999813079834,8,8,0,8,0.10250000000000009,0.07500000000000001,0.6791599988937378,0.1166667565703392,0.9954468607902527,0.0,554,28,ok
|
||||||
|
-iRBcNs9oI8/7,-iRBcNs9oI8,7,cc7b1c7a06ec41d72007b5c42fb8036f30b66d5cf76e3f093160c22fc5e38081,4.104000091552734,9,9,0,9,0.10222222222222223,0.02666666666666669,0.7135127186775208,0.18537555634975433,0.995795726776123,0.0,401,21,ok
|
||||||
|
-iRBcNs9oI8/6,-iRBcNs9oI8,6,030b0b8d8d4ae47799432ab71b66fcc881f077c5b0feef892c1f146dc0a592e3,2.9030001163482666,5,5,0,5,0.03999999999999999,0.04399999999999998,0.7227271199226379,0.12799954414367676,0.9973600506782532,0.0,283,15,ok
|
||||||
|
-iRBcNs9oI8/9,-iRBcNs9oI8,9,352bdcc73d3622d6abac1b09f2039b211d15b9bf7acee66f0b07783c9b7a4c74,3.4159998893737793,5,5,0,5,0.20800000000000002,0.016000000000000014,0.643750011920929,0.14118166267871857,0.993934154510498,0.0,338,17,ok
|
||||||
|
-iRBcNs9oI8/8,-iRBcNs9oI8,8,72a49e3da2862addcde033bd2fbae957d086533664ca4dcbccd3a851bcd50412,8.093000411987305,21,21,0,21,0.23238095238095247,0.2228571428571429,0.6192578673362732,0.18252728879451752,0.9947929382324219,0.0,803,41,ok
|
||||||
|
-lzEya4AM_4/5,-lzEya4AM_4,5,3b74de8d0e66fd754594f05985e5af3480adcc03555f79225d3b727cb4bed043,7.675000190734863,19,19,0,19,0.12000000000000004,0.09473684210526319,0.7427374720573425,0.2366195172071457,0.9901386499404907,1.0,757,64,ok
|
||||||
|
-lzEya4AM_4/6,-lzEya4AM_4,6,ff5da5afb551001d58100ed2471caea51c741046d5396957efbf40082307ca18,14.7919921875,43,43,0,43,0.05023255813953489,0.024651162790697716,0.7318416833877563,0.2524191737174988,0.9907644391059875,1.0,1475,105,ok
|
||||||
|
-mJ2ud6oKI8/1,-mJ2ud6oKI8,1,c8e3acb4ca08679796c4c8cdfcd8c2b3cffc0465015efaffa1eccdb55ff4f45b,6.103000164031982,13,13,0,13,0.733846153846154,0.7399999999999998,0.8174015879631042,0.1573781669139862,1.0,0.0,606,31,ok
|
||||||
|
-mJ2ud6oKI8/2,-mJ2ud6oKI8,2,c41790f8b75f886bd1f5d6d72cbab1985306d0b715e5ae47203f3d0c8a3e7b28,5.730999946594238,11,11,0,11,0.4636363636363635,0.46181818181818185,0.9266790151596069,0.2666672170162201,1.0,0.0,571,29,ok
|
||||||
|
-mJ2ud6oKI8/6,-mJ2ud6oKI8,6,134a3a4ebbc423760133b0da0e90cfe84d75ea85df3968000afd84153c5d828d,2.256999969482422,5,5,0,5,0.11199999999999996,0.06400000000000003,0.6512578129768372,0.15656699240207672,0.9862568378448486,1.0,223,12,ok
|
||||||
|
-mJ2ud6oKI8/9,-mJ2ud6oKI8,9,df3442f3ed8f894b507923e87e426ae9636790a4a502d00c350a215270cc3e47,7.0329999923706055,22,22,0,22,0.11999999999999994,0.09454545454545453,0.5777366757392883,0.25714901089668274,0.9846946001052856,0.9856857061386108,691,35,ok
|
||||||
|
-mJ2ud6oKI8/8,-mJ2ud6oKI8,8,6af6614c9838417a48c4817f4dd5bf1e6669eb5f78f983e9a68adc113fcdc7cc,4.191999912261963,11,11,0,11,0.14363636363636367,0.06363636363636366,0.6529467701911926,0.26233649253845215,0.9846959114074707,1.0,418,21,ok
|
||||||
|
-mqbVkbCndg/0,-mqbVkbCndg,0,cdeba953bba30907814acf72f3e35582dec263eaa25bdd922dad7fafe02e4dda,6.466667175292969,12,12,0,12,0.06166666666666665,0.03333333333333339,0.7074670791625977,0.20000411570072174,0.9905118346214294,1.0,639,53,ok
|
||||||
|
-t217m2on-s/2,-t217m2on-s,2,6713a983112173204502dc5adbb886bc6c38554b4627698611931420c47663a8,5.6860032081604,12,12,0,12,0.21666666666666665,0.2333333333333333,0.6279765367507935,0.16491496562957764,0.9959543347358704,1.0,564,45,ok
|
||||||
|
-t217m2on-s/7,-t217m2on-s,7,41c76b8a733ecb780d56348f55d00f0ba0c07f44f46e93044e62d876710180b7,17.183008193969727,45,45,0,45,0.048444444444444436,0.017777777777777892,0.6867958903312683,0.23202678561210632,0.9928514361381531,1.0,1703,117,ok
|
||||||
|
-tANM6ETl_M/3,-tANM6ETl_M,3,6d3af064c060dad3816a9e1dfa00101faebd8b7da7d8ceea85b0bc0fca70abd6,6.758008003234863,22,22,0,22,0.030909090909090896,0.021818181818181792,0.5859207510948181,0.2757546901702881,0.9936580061912537,0.9898989200592041,661,54,ok
|
||||||
|
-tPCytz4rww/11,-tPCytz4rww,11,6dc09d02467baaef67594ff32cffb01f8ed1565cda4c5abd161b8261a51e2e5d,5.466991901397705,17,17,0,17,0.09764705882352945,0.08588235294117644,0.6753751635551453,0.28518491983413696,0.9890678524971008,1.0,538,43,ok
|
||||||
|
-tPCytz4rww/10,-tPCytz4rww,10,758183192cbd5f0f20823d26ece3ddee40dec0aadec35247b2495510a53f52f8,4.788997173309326,12,12,0,12,0.04499999999999999,0.00833333333333334,0.5449775457382202,0.225633442401886,0.9894654750823975,1.0,468,35,ok
|
||||||
|
-tPCytz4rww/12,-tPCytz4rww,12,e3c7f9cd67ad2d20fef997696dc278361f97f027db58762e22104e5720d509d5,11.863997459411621,26,26,0,26,0.14461538461538456,0.08769230769230774,0.654129683971405,0.22372934222221375,0.9875938892364502,1.0,1174,91,ok
|
||||||
|
-tPCytz4rww/16,-tPCytz4rww,16,8b3f82f9628792eebaf8c227becc2ebeba59bb8abb668bcca86e07fb1a992a93,7.205989837646484,16,16,0,16,0.05125000000000004,0.013749999999999984,0.4489554166793823,0.17846456170082092,0.988937258720398,1.0,714,59,ok
|
||||||
|
-tPCytz4rww/18,-tPCytz4rww,18,eece70451e9c37a4b5f375bf646f66e09113fe3467f35bf1720cc14d986c1776,6.644987106323242,16,16,0,16,0.16375,0.1275,0.6719247698783875,0.22222484648227692,0.9872123599052429,1.0,657,54,ok
|
||||||
|
-vxjVxOeScU/4,-vxjVxOeScU,4,00282df9a314394f19d616d561bd1916d1d30dca4c814dfb43630ca164f4c719,9.127017974853516,21,21,0,21,0.06380952380952383,0.037142857142857144,0.6465563774108887,0.2219761461019516,0.9960808753967285,1.0,904,77,ok
|
||||||
|
-wMB_hJL-3o/7,-wMB_hJL-3o,7,04a73d73fd150b07edab8e652fd14c9cb4a0f6e017ea97b17e8146e377b9fae7,6.044010162353516,22,22,0,22,0.012727272727272738,0.021818181818181816,0.658498227596283,0.31481292843818665,0.9895980954170227,0.9836065769195557,598,48,ok
|
||||||
|
-wny0OAz3g8/1,-wny0OAz3g8,1,c5535859f129ce04da5f5b5104d02d4168057b9d947a5349d423278ad4f3d8e8,6.844009876251221,23,22,1,22,0.013636363636363648,0.011818181818181823,0.62252277135849,0.3443901240825653,0.995954155921936,0.9950981140136719,678,56,ok
|
||||||
|
-wny0OAz3g8/0,-wny0OAz3g8,0,1cc21b4432e2f4b592284ee40558a18acc7d208019117aebdb2f6a1b67ea198d,3.3333330154418945,9,9,0,9,0.32000000000000006,0.2866666666666667,0.6052471995353699,0.2736770212650299,0.9955413937568665,0.9999999403953552,328,21,ok
|
||||||
|
-wny0OAz3g8/3,-wny0OAz3g8,3,ea917c506bc9fa39460f2b722f7c416f646312bb17cd74baa618905b54837bce,6.146028995513916,20,20,0,20,0.022999999999999986,0.02399999999999995,0.7014381289482117,0.3213139474391937,0.9980819821357727,0.9895832538604736,610,49,ok
|
||||||
|
-wny0OAz3g8/2,-wny0OAz3g8,2,1825fb37db2e914fb616cd80aebd613afd873eecb7edd4d0b53158bfda0e522a,6.686978816986084,23,23,0,23,0.02869565217391302,0.026086956521739167,0.7553136348724365,0.33939871191978455,0.9957627058029175,0.9848485589027405,653,54,ok
|
||||||
|
-wny0OAz3g8/5,-wny0OAz3g8,5,6c6a8a4cba654b25190ab1793bfae81c3d2159ecd6cbc02d3fcbd598543acce0,6.9119791984558105,20,20,0,20,0.004999999999999982,0.004999999999999982,0.6063091158866882,0.3181923031806946,0.9961636066436768,0.9895831346511841,675,56,ok
|
||||||
|
-wny0OAz3g8/7,-wny0OAz3g8,7,cbf9cac1048ce191aef781d0c4563eb97f9f3ddd52eb1f9551fa78d31d70e7cd,8.06796932220459,29,29,0,29,0.0186206896551724,0.017241379310344838,0.6700748205184937,0.3224986493587494,0.9956274628639221,1.0,792,68,ok
|
||||||
|
-wny0OAz3g8/9,-wny0OAz3g8,9,b76f70fe783ee76bf1b637137226c98e67d13a6b15ec725894175b36cd76a136,6.2130208015441895,20,20,0,20,0.025999999999999975,0.027000000000000024,0.6688015460968018,0.3180274963378906,0.9969837069511414,0.9672130942344666,605,49,ok
|
||||||
|
-571d8cVauQ/0,-571d8cVauQ,0,defc62e1530d9b0cd27f2acfe042b71dd84d7e8c85f1df261f04912091976088,5.066667079925537,12,12,0,12,0.07500000000000001,0.03000000000000001,0.6978414058685303,0.2360049933195114,0.9903978705406189,1.0,500,39,ok
|
||||||
|
-571d8cVauQ/5,-571d8cVauQ,5,e304c03fac1f477823fe0926c4fb421b9c4e4091f85718b729afe856c24a7d3e,15.272981643676758,43,43,0,43,0.14604651162790694,0.11534883720930222,0.7017207741737366,0.23815667629241943,0.9918006062507629,1.0,1512,108,ok
|
||||||
|
-I_e4mIh0yE/1,-I_e4mIh0yE,1,8453a966d4bd1f0179eb86a35cd34c0864a3a3bc2fa86f37783060cc0d14f5df,7.622004985809326,18,18,0,18,0.061111111111111116,0.02333333333333332,0.6546477675437927,0.2152525633573532,0.985649049282074,1.0,745,63,ok
|
||||||
|
-I_e4mIh0yE/3,-I_e4mIh0yE,3,a015f02ae51d74c42ad27e2a0831263a303f0658ac8d984819e27ad61342b23e,9.163021087646484,22,22,0,22,0.13,0.11545454545454543,0.6251755356788635,0.24003465473651886,0.9852646589279175,1.0,900,77,ok
|
||||||
|
-UacrmKiTn4/10,-UacrmKiTn4,10,00312e0602dbeac970de5d7dd0e8a970e514fe45d7fe6625b4f454b299e36360,7.361979007720947,18,18,0,18,0.04555555555555555,0.07222222222222224,0.47940683364868164,0.20824021100997925,0.9955651164054871,1.0,725,60,ok
|
||||||
|
-UacrmKiTn4/4,-UacrmKiTn4,4,a7f2ab0d6ee2e2a2d5cd4239f03e6cac67c618ddb744891323df3ff87e29af84,4.741015911102295,14,14,0,14,0.06714285714285716,0.09857142857142863,0.7214669585227966,0.246806800365448,0.9926760792732239,1.0,469,35,ok
|
||||||
|
-hnBHBN8p5A/7,-hnBHBN8p5A,7,0ec6fb76226315d3fafbb483e3500e60cc8a6729617b10e4b4b76163595080fb,7.372000217437744,13,13,0,13,0.01846153846153847,0.021538461538461593,0.657243549823761,0.1729772984981537,0.9905768632888794,0.0,731,37,ok
|
||||||
|
-hnBHBN8p5A/6,-hnBHBN8p5A,6,997f6808ac3973ddd71800e91c8e10755c10b5037186d84348a79e00d672fdc7,6.401000022888184,13,13,0,13,0.04307692307692312,0.012307692307692353,0.7618876099586487,0.22187285125255585,0.9932082295417786,0.0,634,32,ok
|
||||||
|
-qDkUB0GgYY/6,-qDkUB0GgYY,6,ab19bf4a761d3919b7113b1104ccaadbbf0c4e7b45694904d094280a2c329c8b,4.677995204925537,16,16,0,16,0.033749999999999974,0.02250000000000002,0.6504627466201782,0.24682332575321198,0.9946085214614868,1.0,462,35,ok
|
||||||
|
-uywlfIYOS8/4,-uywlfIYOS8,4,fde46f5222cf5cfec7af9c93214c972ce2d686b0bf9c90a291021601968986f9,6.044987201690674,18,18,0,18,0.0433333333333333,0.04333333333333334,0.6908547878265381,0.25089454650878906,0.9957761168479919,1.0,589,48,ok
|
||||||
|
-6rXp3zJ3kc/8,-6rXp3zJ3kc,8,59b77a0f873b7fa95e7327b93d64bb5a4bdae84039321988ef70bba8e63ee1ae,12.805012702941895,34,34,0,34,0.06470588235294118,0.034705882352941086,0.63474440574646,0.22114422917366028,0.9899758696556091,1.0,1274,95,ok
|
||||||
|
-9y-fZ3swSY/0,-9y-fZ3swSY,0,29a1fd181293cdbe24e50edcad51e3cc0f27a5824d840bcc39d453b272fb6e3c,6.8333330154418945,22,22,0,22,0.023636363636363688,0.019999999999999997,0.5440928339958191,0.2264895737171173,0.9948647022247314,1.0,676,57,ok
|
||||||
|
-9y-fZ3swSY/4,-9y-fZ3swSY,4,fe04a26710e85c58bdaad5a48618c7e0f742d82b0d920eb4a046c8f14a6921dc,2.8210289478302,9,9,0,9,0.1711111111111111,0.15999999999999998,0.5274830460548401,0.1851968914270401,0.9926167726516724,1.0,267,15,ok
|
||||||
|
-9y-fZ3swSY/8,-9y-fZ3swSY,8,6140d031e614b62a5d8df73346fef2417c28cee3501310aa95aa40d358b4b256,4.988996982574463,12,12,0,12,0.07333333333333339,0.05333333333333331,0.562736988067627,0.18345493078231812,0.9931800365447998,0.9487179517745972,492,37,ok
|
||||||
|
-AUZQgSxyPQ/2,-AUZQgSxyPQ,2,e09034b30bb4f8fb9f9c2568afc20f2a764e327fe6ef4e13f21804eaac51bb7b,22.511003494262695,41,41,0,41,0.6434146341463415,0.6102439024390245,0.6386448740959167,0.19374850392341614,0.9917553663253784,1.0,2241,144,ok
|
||||||
|
-HeZS2-Prhc/2,-HeZS2-Prhc,2,edb0fa22fdfbd990c8174764af78670bd2ecad799240d35d96a1124a3cd9dc0c,8.261979103088379,16,16,0,16,0.14125000000000004,0.08125000000000016,0.6830984950065613,0.15250585973262787,0.9933875799179077,1.0,816,70,ok
|
||||||
|
-MeTTeMJBNc/0,-MeTTeMJBNc,0,54e2097f7bdb455d243e72254e5af7ed16160739f2f3cfd22f57d2729ec1acbe,9.300000190734863,21,21,0,21,0.03714285714285711,0.026666666666666637,0.6400595903396606,0.22926528751850128,0.9939717650413513,1.0,922,78,ok
|
||||||
|
-MeTTeMJBNc/13,-MeTTeMJBNc,13,80c5d996e5d7b61c58e1fe32d706efcaf80c8c8cddb5c37a5c66d527786c2a44,5.430013179779053,14,14,0,14,0.042857142857142864,0.00714285714285715,0.6661255359649658,0.2075468748807907,0.9918515682220459,1.0,526,41,ok
|
||||||
|
-MeTTeMJBNc/7,-MeTTeMJBNc,7,40716dc7402fd6398038eded58ce939ec76be6afbeea346b968bf4c427088f3a,10.51699161529541,31,31,0,31,0.03225806451612902,0.02387096774193552,0.7313360571861267,0.269477516412735,0.9934103488922119,1.0,1042,84,ok
|
||||||
|
-RfYyzHpjk4/11,-RfYyzHpjk4,11,96e481cd732001239639c8cb4e7f94e570925e5a15940920ffab0cec1f13f904,5.238996982574463,17,17,0,17,0.018823529411764742,0.025882352941176467,0.6960803866386414,0.27842891216278076,0.9958056807518005,1.0,508,40,ok
|
||||||
|
-RfYyzHpjk4/8,-RfYyzHpjk4,8,ed8c672c532b8c7d1c830c82d6e2c14fcb3bbd3dc1a86670a03a6c8e83c02f3c,4.588996887207031,17,17,0,17,0.04588235294117646,0.005882352941176463,0.684955358505249,0.27527740597724915,0.994469404220581,1.0,454,33,ok
|
||||||
|
-RfYyzHpjk4/2,-RfYyzHpjk4,2,5e428ee22de58c9d6561b9b4353bef9dfb811a5841d9ab07b5bc453d697ad17f,5.516016006469727,25,25,0,25,0.12960000000000005,0.12880000000000008,0.7175998091697693,0.3444570004940033,0.9948742389678955,1.0,540,43,ok
|
||||||
|
-UUCSKoHeMA/0,-UUCSKoHeMA,0,3f9792888ec8e969f61cfb6cb4de7bdf7ef8944afe0a9d2a9d13586e1f13b898,7.800000190734863,18,18,0,18,0.23000000000000004,0.20999999999999996,0.666740894317627,0.22424933314323425,0.9943835139274597,1.0,774,66,ok
|
||||||
|
-ri04Z7vwnc/0,-ri04Z7vwnc,0,93021c70c8fad20ad3fded80e7fc790a2f9c96835de58a3b694c6906de42684b,2.806999921798706,8,8,0,8,0.35999999999999993,0.3525,0.7826409935951233,0.25714799761772156,0.9982975125312805,0.0,277,14,ok
|
||||||
|
-ri04Z7vwnc/2,-ri04Z7vwnc,2,f5993ce3a6ee56298462c08692a96f78c7d914d7264257ed401de7d6e17fd148,6.0269999504089355,16,16,0,16,0.23374999999999996,0.22875,0.7431082725524902,0.26666611433029175,0.9971521496772766,0.9906440377235413,592,30,ok
|
||||||
|
-ri04Z7vwnc/5,-ri04Z7vwnc,5,2a5e4319bf592a18c8eb5117f7c7ec30c6c79528f9eac9fe4fb85bf70f277736,3.5450000762939453,9,8,1,8,0.21249999999999986,0.23999999999999994,0.721091628074646,0.21714359521865845,0.9974266886711121,1.0,344,18,ok
|
||||||
|
-s9qJ7ATP7w/1,-s9qJ7ATP7w,1,a33e8f82a185bcb7648b523927308cdc2142e3076479ba2b45cd0e828636fc09,4.7919921875,11,11,0,11,0.01636363636363641,0.021818181818181816,0.7055363059043884,0.16956037282943726,0.9860281348228455,1.0,465,35,ok
|
||||||
|
-s9qJ7ATP7w/0,-s9qJ7ATP7w,0,f2324dbc6debe0d1e4254e6f943523373e3cbb30d8d0555413ce3a7b370a4af5,6.800000190734863,24,24,0,24,0.056666666666666664,0.029999999999999957,0.6351987719535828,0.2500007152557373,0.9885490536689758,0.984375,672,56,ok
|
||||||
|
-s9qJ7ATP7w/5,-s9qJ7ATP7w,5,bd35eb8de64ce1f8b27921294598f7cf5bb971255738f51d06458c7096c58eae,8.561002731323242,19,19,0,19,0.25578947368421046,0.26210526315789456,0.5150972604751587,0.16428017616271973,0.9896790981292725,1.0,840,72,ok
|
||||||
|
-s9qJ7ATP7w/4,-s9qJ7ATP7w,4,76e0672ffce5857a789b41fb10606402f7da58c33579b1032a8282f86fb35bc9,4.427018165588379,9,9,0,9,0.07555555555555557,0.024444444444444453,0.6215080618858337,0.20000608265399933,0.9892621040344238,0.96875,425,31,ok
|
||||||
|
-s9qJ7ATP7w/7,-s9qJ7ATP7w,7,5301f4403cf7cd299cc525d2f443392ac70fd70c3c9cd4a09406f2ba4df6734d,6.966015815734863,28,28,0,28,0.029285714285714304,0.02428571428571429,0.661829948425293,0.29855138063430786,0.9849807620048523,0.9496123194694519,684,57,ok
|
||||||
|
-s9qJ7ATP7w/6,-s9qJ7ATP7w,6,40b684fd1b4559f0f82e72f9b47f4ce7c5e15059fe8d0d8ab112637cd3386451,2.4749999046325684,5,5,0,5,0.21599999999999997,0.16400000000000006,0.8113580942153931,0.14165958762168884,0.9844902753829956,1.0,232,14,ok
|
||||||
|
-s9qJ7ATP7w/8,-s9qJ7ATP7w,8,1de4d91bd9d9fb508a3779658747da88460db0bdb1beefe7ea5b2c77edfd2239,3.2949869632720947,13,12,1,12,0.07166666666666661,0.06166666666666667,0.6465047001838684,0.2687399089336395,0.9880942702293396,0.9731181859970093,319,20,ok
|
||||||
|
-yRb-Jum7EQ/1,-yRb-Jum7EQ,1,2f16da1baa8660e5bfc608f5237cdd59725d16e57a19bc02c4556cd2c7129a46,29.288021087646484,49,49,0,49,0.2922448979591836,0.2689795918367347,0.6873865723609924,0.18630558252334595,0.9836021065711975,0.999393880367279,2911,178,ok
|
||||||
|
-yRb-Jum7EQ/5,-yRb-Jum7EQ,5,6d6145293cf84d1fc8324dfa8d39465e2d7fc07affb67ff68deeabf6e3c410d9,13.169010162353516,25,25,0,25,0.14880000000000007,0.18799999999999997,0.7192046046257019,0.19541765749454498,0.9867846369743347,1.0,1305,97,ok
|
||||||
|
-yRb-Jum7EQ/6,-yRb-Jum7EQ,6,b51eb9ad06de804934cb6a315de591ba9ff7f8aa71bffff7f8001ab29945d989,11.241994857788086,19,19,0,19,0.09578947368421038,0.15263157894736842,0.6869907975196838,0.16814078390598297,0.9854583740234375,1.0,1125,87,ok
|
||||||
|
@@ -0,0 +1,101 @@
|
|||||||
|
sample_id,video_id,fold,split
|
||||||
|
-3g5yACwYnA/13,-3g5yACwYnA,1,valid_oof
|
||||||
|
-3g5yACwYnA/2,-3g5yACwYnA,1,valid_oof
|
||||||
|
-3g5yACwYnA/3,-3g5yACwYnA,1,valid_oof
|
||||||
|
-3g5yACwYnA/9,-3g5yACwYnA,1,valid_oof
|
||||||
|
-3nNcZdcdvU/5,-3nNcZdcdvU,5,valid_oof
|
||||||
|
-571d8cVauQ/0,-571d8cVauQ,2,valid_oof
|
||||||
|
-571d8cVauQ/5,-571d8cVauQ,2,valid_oof
|
||||||
|
-6rXp3zJ3kc/8,-6rXp3zJ3kc,4,valid_oof
|
||||||
|
-9y-fZ3swSY/0,-9y-fZ3swSY,1,valid_oof
|
||||||
|
-9y-fZ3swSY/4,-9y-fZ3swSY,1,valid_oof
|
||||||
|
-9y-fZ3swSY/8,-9y-fZ3swSY,1,valid_oof
|
||||||
|
-AUZQgSxyPQ/2,-AUZQgSxyPQ,3,valid_oof
|
||||||
|
-HeZS2-Prhc/2,-HeZS2-Prhc,2,valid_oof
|
||||||
|
-HwX2H8Z4hY/2,-HwX2H8Z4hY,3,valid_oof
|
||||||
|
-HwX2H8Z4hY/5,-HwX2H8Z4hY,3,valid_oof
|
||||||
|
-HwX2H8Z4hY/6,-HwX2H8Z4hY,3,valid_oof
|
||||||
|
-HwX2H8Z4hY/9,-HwX2H8Z4hY,3,valid_oof
|
||||||
|
-I_e4mIh0yE/1,-I_e4mIh0yE,1,valid_oof
|
||||||
|
-I_e4mIh0yE/3,-I_e4mIh0yE,1,valid_oof
|
||||||
|
-MeTTeMJBNc/0,-MeTTeMJBNc,5,valid_oof
|
||||||
|
-MeTTeMJBNc/13,-MeTTeMJBNc,5,valid_oof
|
||||||
|
-MeTTeMJBNc/7,-MeTTeMJBNc,5,valid_oof
|
||||||
|
-NFrJFQijFE/1,-NFrJFQijFE,5,valid_oof
|
||||||
|
-NFrJFQijFE/2,-NFrJFQijFE,5,valid_oof
|
||||||
|
-RfYyzHpjk4/11,-RfYyzHpjk4,3,valid_oof
|
||||||
|
-RfYyzHpjk4/2,-RfYyzHpjk4,3,valid_oof
|
||||||
|
-RfYyzHpjk4/8,-RfYyzHpjk4,3,valid_oof
|
||||||
|
-THoVjtIkeU/12,-THoVjtIkeU,2,valid_oof
|
||||||
|
-THoVjtIkeU/2,-THoVjtIkeU,2,valid_oof
|
||||||
|
-THoVjtIkeU/6,-THoVjtIkeU,2,valid_oof
|
||||||
|
-UUCSKoHeMA/0,-UUCSKoHeMA,1,valid_oof
|
||||||
|
-UacrmKiTn4/10,-UacrmKiTn4,4,valid_oof
|
||||||
|
-UacrmKiTn4/4,-UacrmKiTn4,4,valid_oof
|
||||||
|
-UuX1xuaiiE/0,-UuX1xuaiiE,2,valid_oof
|
||||||
|
-UuX1xuaiiE/1,-UuX1xuaiiE,2,valid_oof
|
||||||
|
-UuX1xuaiiE/3,-UuX1xuaiiE,2,valid_oof
|
||||||
|
-UuX1xuaiiE/6,-UuX1xuaiiE,2,valid_oof
|
||||||
|
-a55Q6RWvTA/3,-a55Q6RWvTA,5,valid_oof
|
||||||
|
-aNfi7CP8vM/7,-aNfi7CP8vM,4,valid_oof
|
||||||
|
-aqamKhZ1Ec/0,-aqamKhZ1Ec,3,valid_oof
|
||||||
|
-dxfTGcXJoc/0,-dxfTGcXJoc,5,valid_oof
|
||||||
|
-dxfTGcXJoc/1,-dxfTGcXJoc,5,valid_oof
|
||||||
|
-dxfTGcXJoc/2,-dxfTGcXJoc,5,valid_oof
|
||||||
|
-dxfTGcXJoc/6,-dxfTGcXJoc,5,valid_oof
|
||||||
|
-egA8-b7-3M/1,-egA8-b7-3M,1,valid_oof
|
||||||
|
-egA8-b7-3M/13,-egA8-b7-3M,1,valid_oof
|
||||||
|
-egA8-b7-3M/16,-egA8-b7-3M,1,valid_oof
|
||||||
|
-egA8-b7-3M/17,-egA8-b7-3M,1,valid_oof
|
||||||
|
-egA8-b7-3M/18,-egA8-b7-3M,1,valid_oof
|
||||||
|
-egA8-b7-3M/20,-egA8-b7-3M,1,valid_oof
|
||||||
|
-egA8-b7-3M/26,-egA8-b7-3M,1,valid_oof
|
||||||
|
-egA8-b7-3M/6,-egA8-b7-3M,1,valid_oof
|
||||||
|
-egA8-b7-3M/9,-egA8-b7-3M,1,valid_oof
|
||||||
|
-hnBHBN8p5A/6,-hnBHBN8p5A,3,valid_oof
|
||||||
|
-hnBHBN8p5A/7,-hnBHBN8p5A,3,valid_oof
|
||||||
|
-iRBcNs9oI8/3,-iRBcNs9oI8,4,valid_oof
|
||||||
|
-iRBcNs9oI8/6,-iRBcNs9oI8,4,valid_oof
|
||||||
|
-iRBcNs9oI8/7,-iRBcNs9oI8,4,valid_oof
|
||||||
|
-iRBcNs9oI8/8,-iRBcNs9oI8,4,valid_oof
|
||||||
|
-iRBcNs9oI8/9,-iRBcNs9oI8,4,valid_oof
|
||||||
|
-lzEya4AM_4/5,-lzEya4AM_4,2,valid_oof
|
||||||
|
-lzEya4AM_4/6,-lzEya4AM_4,2,valid_oof
|
||||||
|
-mJ2ud6oKI8/1,-mJ2ud6oKI8,5,valid_oof
|
||||||
|
-mJ2ud6oKI8/2,-mJ2ud6oKI8,5,valid_oof
|
||||||
|
-mJ2ud6oKI8/6,-mJ2ud6oKI8,5,valid_oof
|
||||||
|
-mJ2ud6oKI8/8,-mJ2ud6oKI8,5,valid_oof
|
||||||
|
-mJ2ud6oKI8/9,-mJ2ud6oKI8,5,valid_oof
|
||||||
|
-mqbVkbCndg/0,-mqbVkbCndg,2,valid_oof
|
||||||
|
-qDkUB0GgYY/6,-qDkUB0GgYY,1,valid_oof
|
||||||
|
-ri04Z7vwnc/0,-ri04Z7vwnc,4,valid_oof
|
||||||
|
-ri04Z7vwnc/2,-ri04Z7vwnc,4,valid_oof
|
||||||
|
-ri04Z7vwnc/5,-ri04Z7vwnc,4,valid_oof
|
||||||
|
-s9qJ7ATP7w/0,-s9qJ7ATP7w,3,valid_oof
|
||||||
|
-s9qJ7ATP7w/1,-s9qJ7ATP7w,3,valid_oof
|
||||||
|
-s9qJ7ATP7w/4,-s9qJ7ATP7w,3,valid_oof
|
||||||
|
-s9qJ7ATP7w/5,-s9qJ7ATP7w,3,valid_oof
|
||||||
|
-s9qJ7ATP7w/6,-s9qJ7ATP7w,3,valid_oof
|
||||||
|
-s9qJ7ATP7w/7,-s9qJ7ATP7w,3,valid_oof
|
||||||
|
-s9qJ7ATP7w/8,-s9qJ7ATP7w,3,valid_oof
|
||||||
|
-t217m2on-s/2,-t217m2on-s,4,valid_oof
|
||||||
|
-t217m2on-s/7,-t217m2on-s,4,valid_oof
|
||||||
|
-tANM6ETl_M/3,-tANM6ETl_M,5,valid_oof
|
||||||
|
-tPCytz4rww/10,-tPCytz4rww,4,valid_oof
|
||||||
|
-tPCytz4rww/11,-tPCytz4rww,4,valid_oof
|
||||||
|
-tPCytz4rww/12,-tPCytz4rww,4,valid_oof
|
||||||
|
-tPCytz4rww/16,-tPCytz4rww,4,valid_oof
|
||||||
|
-tPCytz4rww/18,-tPCytz4rww,4,valid_oof
|
||||||
|
-uywlfIYOS8/4,-uywlfIYOS8,4,valid_oof
|
||||||
|
-vxjVxOeScU/4,-vxjVxOeScU,3,valid_oof
|
||||||
|
-wMB_hJL-3o/7,-wMB_hJL-3o,3,valid_oof
|
||||||
|
-wny0OAz3g8/0,-wny0OAz3g8,2,valid_oof
|
||||||
|
-wny0OAz3g8/1,-wny0OAz3g8,2,valid_oof
|
||||||
|
-wny0OAz3g8/2,-wny0OAz3g8,2,valid_oof
|
||||||
|
-wny0OAz3g8/3,-wny0OAz3g8,2,valid_oof
|
||||||
|
-wny0OAz3g8/5,-wny0OAz3g8,2,valid_oof
|
||||||
|
-wny0OAz3g8/7,-wny0OAz3g8,2,valid_oof
|
||||||
|
-wny0OAz3g8/9,-wny0OAz3g8,2,valid_oof
|
||||||
|
-yRb-Jum7EQ/1,-yRb-Jum7EQ,5,valid_oof
|
||||||
|
-yRb-Jum7EQ/5,-yRb-Jum7EQ,5,valid_oof
|
||||||
|
-yRb-Jum7EQ/6,-yRb-Jum7EQ,5,valid_oof
|
||||||
|
Binary file not shown.
|
After Width: | Height: | Size: 1.0 MiB |
File diff suppressed because it is too large
Load Diff
Binary file not shown.
Binary file not shown.
Reference in New Issue
Block a user