整理 Q1-Q3 实验代码与结果
This commit is contained in:
@@ -270,6 +270,7 @@ stay fixed. Emotion labels are not used.
|
||||
cd Q1
|
||||
uv run python -m q1.tsfa_experiment --device cuda --seed 42
|
||||
uv run python -m q1.tsfa_alignment_only_eval --device cuda --seed 42
|
||||
uv run python -m q1.tsfa_emotion_probe --device cuda --seed 42
|
||||
```
|
||||
|
||||
The first command fits one semantic branch and equal-seed correspondence probes
|
||||
@@ -282,5 +283,101 @@ paired video-group confidence intervals, monotonicity/span/entropy diagnostics,
|
||||
plots, manifests, and checkpoints are written to `outputs/tsfa/`; the small
|
||||
`report_bundle/` contains shareable
|
||||
summaries and figures. See [RESULTS.md](RESULTS.md) for the measured tradeoff
|
||||
and interpretation limits.
|
||||
and interpretation limits. The third command evaluates the same frozen
|
||||
representations with the grouped, five-segment logistic-regression and Ridge
|
||||
probes; emotion labels are used only by these downstream probes.
|
||||
|
||||
With the math OOF predictions and fold assignment CSV available at their
|
||||
workspace paths, compare TSFA against B0–B4 on the identical held-out samples
|
||||
and bootstrap by source video:
|
||||
|
||||
```bash
|
||||
uv run python -m q1.compare_emotion_probes --bootstrap-repeats 2000 --seed 42
|
||||
```
|
||||
|
||||
The local Audio–Vision edge and modality-private residual experiment runs a
|
||||
fixed 2×2 ablation with the same five grouped folds, temporal windows,
|
||||
self-supervised objective, and emotion probes:
|
||||
|
||||
```bash
|
||||
uv run python -m q1.tsfa_av_private_ablation --device cuda --seed 42
|
||||
uv run python -m q1.compare_emotion_probes --tsfa-predictions outputs/tsfa_av_private_ablation/predictions.csv --candidate-method TSFA-T+Private --output-dir outputs/tsfa_av_private_ablation --output-stem private_vs_math --bootstrap-repeats 2000 --seed 42
|
||||
```
|
||||
|
||||
The ablation files and figure are in `outputs/tsfa_av_private_ablation/`.
|
||||
|
||||
## TSFA-SPR shared/private factorization
|
||||
|
||||
Run the first grouped five-fold experiment with M4 temporal checkpoints frozen:
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.tsfa_shared_private --device cuda --seed 42
|
||||
```
|
||||
|
||||
The experiment pools each fold-standardized original text, audio, and vision
|
||||
source stream with the matching frozen `M4_sourceTime` attention weights. A
|
||||
shared MLP and three independent private MLPs learn from same-slot cross-modal
|
||||
InfoNCE, within-modality orthogonality, and per-modality reconstruction. No
|
||||
emotion labels enter representation training. The run also fits a train-fold
|
||||
PCA control on raw private slots to match the 256-dimensional-per-slot SPR
|
||||
control, then evaluates all representations with the existing grouped
|
||||
LogisticRegression/Ridge probes. Outputs include the fold checkpoints,
|
||||
training history, OOF predictions, grouped bootstrap contrasts, source and
|
||||
reconstruction diagnostics, and figures under
|
||||
`outputs/tsfa_shared_private/`.
|
||||
|
||||
If you only need to regenerate the derived source contrasts, assessment, and
|
||||
plots from an existing run without retraining, use:
|
||||
|
||||
```bash
|
||||
uv run python -m q1.tsfa_shared_private --finalize-existing
|
||||
```
|
||||
|
||||
For the first seed's interpretation, diagnostic caveats, and next-step
|
||||
decision, see [RESULTS.md](RESULTS.md#tsfa-spr时序共享私有分解) and
|
||||
[`outputs/tsfa_shared_private/README.md`](outputs/tsfa_shared_private/README.md).
|
||||
|
||||
## Shared information definition comparison
|
||||
|
||||
Compare three concrete definitions of shared information using the same
|
||||
fold-standardized original source features pooled by the frozen M4 temporal
|
||||
checkpoints: Similarity (reuse the existing SPR weights), Correlation
|
||||
(regularized MAXVAR/GCCA), and Predictability (cross-modal Ridge prediction
|
||||
and its private residual). The Similarity network is loaded for inference and
|
||||
is not retrained. Emotion labels are used only by the identical outer-fold
|
||||
emotion probes.
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.shared_definition_comparison --device cuda --seed 42
|
||||
```
|
||||
|
||||
The run reuses the fixed five-fold `video_id` splits, compares same-slot,
|
||||
shifted, and within-video-shuffled predictive controls, and reports shared
|
||||
variance, residual predictability, dimension-matched 1,280D emotion probes,
|
||||
private-source ablations, and paired video-group bootstrap contrasts against
|
||||
RawPrivate-PCA, Similarity-SPR, TSFA+RawPrivate, and math B0/B4. The math
|
||||
artifacts are read-only inputs. Tables, figures, fold assignments, and the run
|
||||
manifest are written to `outputs/shared_definition_comparison/`; measured
|
||||
results and limitations are in [RESULTS.md](RESULTS.md#shared-information-definition-similarity-vs-correlation-vs-predictability).
|
||||
|
||||
## Shared time coordinate with modality-private content
|
||||
|
||||
This follow-up uses the corrected definition: M4 supplies a frozen 50-slot
|
||||
temporal coordinate, while BERT text, eGeMAPS audio, and DeiT vision retain
|
||||
modality-specific content. Optional same-slot product/difference features are
|
||||
evaluated only as task-specific interactions; they do not assert shared
|
||||
semantics. Feature standardization, PCA, and emotion probes are fitted inside
|
||||
each training fold. The math outputs are read-only historical controls.
|
||||
|
||||
```bash
|
||||
cd Q1
|
||||
uv run python -m q1.shared_time_private_content --device cuda --seed 42 --bootstrap-draws 2000
|
||||
```
|
||||
|
||||
The run writes 19 new representation OOF probes, prior-result controls,
|
||||
video-group bootstrap contrasts, temporal monitoring, fold/PCA audits, and nine
|
||||
figures to `outputs/shared_time_private_content/`. See the dedicated
|
||||
[results and seven-question interpretation](RESULTS.md#shared-time-private-content修正版验证).
|
||||
|
||||
|
||||
Reference in New Issue
Block a user