TSFA five-fold result bundle
Seven frozen-alignment comparisons on 100 held-out clips from 37 video IDs. All Python runs use the project uv environment. Read run_manifest.json for the complete protocol and limitations. paired_contrasts.csv reports left minus right with video-ID-cluster bootstrap confidence intervals. The full outputs/tsfa/ directory retains per-clip CSVs and probe checkpoints.
The alignment_only_* summaries pool identical train-fold-standardized raw source features through each method's alignment matrix. They isolate source selection from native model value/output projections. tsfa_temporal_diagnostics_summary.csv reports MVR, span, entropy, and source coverage.
emotion_probe_metrics.csv adds five-fold, video-group-held-out LogisticRegression/Ridge probes on five-segment pooled representations. The classification metrics use fixed three-class Macro-F1; Ridge predictions are clipped to [-3, 3] before reported MAE/Pearson. The CSV also includes fold Macro-F1 mean/sample SD and unclipped regression metrics for diagnosis. These are small-sample downstream probes, not end-to-end emotion model scores.
paired_math_comparison.csv compares the TSFA all-modalities OOF predictions with math B0–B4 using the identical 100 sample IDs and fold assignments. Confidence intervals use paired bootstrap resampling of the 37 source-video groups; the comparison script only reads the math prediction and split CSVs.