Q1 method comparison
This folder contains grouped cross-validation results for M1–M4. M1/M2 are fixed rules; M3/M4 are trained without emotion labels. All learned models and probes use training-fold-only feature normalization, and folds are grouped by video_id.
alignment_metrics.csv reports expected-time trajectories, monotonicity, attention entropy, and width diagnostics. retrieval_probe_metrics.csv uses a separately trained linear projection probe; its grid-index positives are not independent temporal ground truth. reconstruction_probe_metrics.csv reports held-out masked reconstruction error in training-fold standardized feature units. frozen_emotion_probe_metrics.csv is a small-sample downstream utility check.
No human event-time annotation is present, so the results cannot establish direct human alignment accuracy. See run_manifest.json for parameters, seeds, device, and interpretation limits.