Files
modeling_zhaocui/deep_learning/Q3/RESULTS.md
T

4.2 KiB
Raw Blame History

Q3 explanation algorithm selection results

Predictor and validation setup

Q3 reuses the Q2-selected early-concatenation + BiGRU classifier/regressor. Explanations were compared on all 728 held-out Attachment 2 validation clips; the predictor was trained on the official training split. Five-slot groups give 30 possible modality/time regions per clip. The class target is each clip's predicted-class probability; the intensity target is the predicted continuous score.

Explanation comparison

Explainer Target Signed change after deleting top 30% Absolute change after deleting top 30% Error when keeping only top 30% Deletion AUC, 10–50% Runtime for 728 clips
Grouped occlusion Class probability +0.260 0.274 0.024 0.099 0.40 s
Integrated Gradients Class probability +0.258 0.283 0.040 0.093 4.72 s
Random-region control Class probability +0.057 0.077 0.163 0.023 —
Grouped occlusion Intensity −0.132 0.588 0.070 −0.058 0.40 s
Integrated Gradients Intensity −0.060 0.632 0.114 −0.036 4.72 s
Random-region control Intensity −0.031 0.185 0.377 −0.011 —

Grouped occlusion is selected for polarity: it produces a slightly larger signed class-probability drop, lower sufficiency error, higher deletion AUC, and runs about 12 times faster. For intensity, the result is a tradeoff. Integrated Gradients causes a larger prediction change when its top evidence is removed; grouped occlusion better preserves the prediction when only its top regions remain. The displayed intensity regions use grouped occlusion, with Integrated Gradients retained as a directional cross-check. The negative signed intensity changes mean that removing the selected regions raises the predicted score on average; intensity evidence is bidirectional.

Under standardized input noise with σ=0.02, the top-region rank Spearman correlations were 0.9993–0.9999 and top-30% Jaccard overlap was 0.987–0.998 on 120 balanced validation clips. This shows stability to that small perturbation, not stability across retrained models or a different dataset.

Mapping Attachment 4 evidence to video time

All 20 Attachment 4 MP4 files were found, and all 20 BERT token sequences matched the supplied feature token IDs. Q1's hard CTC Viterbi word-time procedure aligned at least one word in every clip; 19 clips had full transcript word coverage, with mean word coverage 99.3%. Text wordpieces and corresponding aligned audio/vision slots inherit the transcript word interval, allowing a selected five-slot region to be shown in clip seconds.

CTC times are weak alignment references, not human event labels. One transcript has partial coverage. The CTC path score is uncalibrated, so it is recorded for review and is not presented as a probability or ground truth. Attachment 4's pkl files themselves do not contain Q1 time_bounds_s; the script computes word times from the supplied video audio and transcript.

Example timeline for clip 01:

Attachment 4 clip 01 explanation timeline

Artifacts

The top-region CSV carries slot indices, transcript words, CTC-derived start/end seconds, uncalibrated alignment quality, prediction outputs, and modality-specific importance. It can be used to inspect individual samples or prepare the Chapter 4 evidence examples.