整理 Q1-Q3 实验代码与结果

This commit is contained in:
2026-09-24 16:25:15 +08:00
parent 0261ecdfba
commit 8f5c2c3be6
247 changed files with 69828 additions and 19 deletions
+98 -1
View File
@@ -270,6 +270,7 @@ stay fixed. Emotion labels are not used.
cd Q1
uv run python -m q1.tsfa_experiment --device cuda --seed 42
uv run python -m q1.tsfa_alignment_only_eval --device cuda --seed 42
uv run python -m q1.tsfa_emotion_probe --device cuda --seed 42
```
The first command fits one semantic branch and equal-seed correspondence probes
@@ -282,5 +283,101 @@ paired video-group confidence intervals, monotonicity/span/entropy diagnostics,
plots, manifests, and checkpoints are written to `outputs/tsfa/`; the small
`report_bundle/` contains shareable
summaries and figures. See [RESULTS.md](RESULTS.md) for the measured tradeoff
and interpretation limits.
and interpretation limits. The third command evaluates the same frozen
representations with the grouped, five-segment logistic-regression and Ridge
probes; emotion labels are used only by these downstream probes.
With the math OOF predictions and fold assignment CSV available at their
workspace paths, compare TSFA against B0–B4 on the identical held-out samples
and bootstrap by source video:
```bash
uv run python -m q1.compare_emotion_probes --bootstrap-repeats 2000 --seed 42
```
The local Audio–Vision edge and modality-private residual experiment runs a
fixed 2×2 ablation with the same five grouped folds, temporal windows,
self-supervised objective, and emotion probes:
```bash
uv run python -m q1.tsfa_av_private_ablation --device cuda --seed 42
uv run python -m q1.compare_emotion_probes --tsfa-predictions outputs/tsfa_av_private_ablation/predictions.csv --candidate-method TSFA-T+Private --output-dir outputs/tsfa_av_private_ablation --output-stem private_vs_math --bootstrap-repeats 2000 --seed 42
```
The ablation files and figure are in `outputs/tsfa_av_private_ablation/`.
## TSFA-SPR shared/private factorization
Run the first grouped five-fold experiment with M4 temporal checkpoints frozen:
```bash
cd Q1
uv run python -m q1.tsfa_shared_private --device cuda --seed 42
```
The experiment pools each fold-standardized original text, audio, and vision
source stream with the matching frozen `M4_sourceTime` attention weights. A
shared MLP and three independent private MLPs learn from same-slot cross-modal
InfoNCE, within-modality orthogonality, and per-modality reconstruction. No
emotion labels enter representation training. The run also fits a train-fold
PCA control on raw private slots to match the 256-dimensional-per-slot SPR
control, then evaluates all representations with the existing grouped
LogisticRegression/Ridge probes. Outputs include the fold checkpoints,
training history, OOF predictions, grouped bootstrap contrasts, source and
reconstruction diagnostics, and figures under
`outputs/tsfa_shared_private/`.
If you only need to regenerate the derived source contrasts, assessment, and
plots from an existing run without retraining, use:
```bash
uv run python -m q1.tsfa_shared_private --finalize-existing
```
For the first seed's interpretation, diagnostic caveats, and next-step
decision, see [RESULTS.md](RESULTS.md#tsfa-spr时序共享私有分解) and
[`outputs/tsfa_shared_private/README.md`](outputs/tsfa_shared_private/README.md).
## Shared information definition comparison
Compare three concrete definitions of shared information using the same
fold-standardized original source features pooled by the frozen M4 temporal
checkpoints: Similarity (reuse the existing SPR weights), Correlation
(regularized MAXVAR/GCCA), and Predictability (cross-modal Ridge prediction
and its private residual). The Similarity network is loaded for inference and
is not retrained. Emotion labels are used only by the identical outer-fold
emotion probes.
```bash
cd Q1
uv run python -m q1.shared_definition_comparison --device cuda --seed 42
```
The run reuses the fixed five-fold `video_id` splits, compares same-slot,
shifted, and within-video-shuffled predictive controls, and reports shared
variance, residual predictability, dimension-matched 1,280D emotion probes,
private-source ablations, and paired video-group bootstrap contrasts against
RawPrivate-PCA, Similarity-SPR, TSFA+RawPrivate, and math B0/B4. The math
artifacts are read-only inputs. Tables, figures, fold assignments, and the run
manifest are written to `outputs/shared_definition_comparison/`; measured
results and limitations are in [RESULTS.md](RESULTS.md#shared-information-definition-similarity-vs-correlation-vs-predictability).
## Shared time coordinate with modality-private content
This follow-up uses the corrected definition: M4 supplies a frozen 50-slot
temporal coordinate, while BERT text, eGeMAPS audio, and DeiT vision retain
modality-specific content. Optional same-slot product/difference features are
evaluated only as task-specific interactions; they do not assert shared
semantics. Feature standardization, PCA, and emotion probes are fitted inside
each training fold. The math outputs are read-only historical controls.
```bash
cd Q1
uv run python -m q1.shared_time_private_content --device cuda --seed 42 --bootstrap-draws 2000
```
The run writes 19 new representation OOF probes, prior-result controls,
video-group bootstrap contrasts, temporal monitoring, fold/PCA audits, and nine
figures to `outputs/shared_time_private_content/`. See the dedicated
[results and seven-question interpretation](RESULTS.md#shared-time-private-content修正版验证).