整理 Q1-Q3 实验代码与结果
This commit is contained in:
@@ -0,0 +1,3 @@
|
||||
.venv/
|
||||
__pycache__/
|
||||
*.py[cod]
|
||||
@@ -0,0 +1,18 @@
|
||||
# Q3 explainability algorithm selection
|
||||
|
||||
Q3 reuses the Q2-selected predictor and its train-only robust scaler. It does not train a different predictor just to make an attribution method look better.
|
||||
|
||||
The experiment compares Integrated Gradients with five-slot grouped occlusion on the held-out Attachment 2 validation split. It measures deletion comprehensiveness, sufficiency, local rank stability under small input noise, and runtime. Polarity and continuous intensity explanations are selected separately because the two outputs can depend on different evidence.
|
||||
|
||||
For the 20 Attachment 4 videos, the script reads the aligned feature pickle and matching MP4, predicts polarity/intensity, then applies the Q1 hard CTC Viterbi word-time procedure to each transcript and source audio. Text wordpieces and aligned audio/vision slots inherit the CTC word interval. The output includes CTC quality and validity flags; CTC timestamps are a weak temporal reference, not human event annotations or ground truth.
|
||||
|
||||
## Run
|
||||
|
||||
The environment is the `uv`-managed Q2 environment, which contains the same CUDA PyTorch, Transformers, NumPy, and plotting dependencies:
|
||||
|
||||
```bash
|
||||
cd deep_learning/Q3
|
||||
uv run --project ../Q2 python -m q3.explain_selection
|
||||
```
|
||||
|
||||
Results are written to `deep_learning/Q3/outputs/explanation_selection/`. Q2 model checkpoints and normalization statistics are read from `deep_learning/Q2/outputs/algorithm_selection/`.
|
||||
@@ -0,0 +1,42 @@
|
||||
# Q3 explanation algorithm selection results
|
||||
|
||||
## Predictor and validation setup
|
||||
|
||||
Q3 reuses the Q2-selected early-concatenation + BiGRU classifier/regressor. Explanations were compared on all 728 held-out Attachment 2 validation clips; the predictor was trained on the official training split. Five-slot groups give 30 possible modality/time regions per clip. The class target is each clip's predicted-class probability; the intensity target is the predicted continuous score.
|
||||
|
||||
## Explanation comparison
|
||||
|
||||
| Explainer | Target | Signed change after deleting top 30% | Absolute change after deleting top 30% | Error when keeping only top 30% | Deletion AUC, 10–50% | Runtime for 728 clips |
|
||||
| --- | --- | ---: | ---: | ---: | ---: | ---: |
|
||||
| Grouped occlusion | Class probability | **+0.260** | 0.274 | **0.024** | **0.099** | **0.40 s** |
|
||||
| Integrated Gradients | Class probability | +0.258 | **0.283** | 0.040 | 0.093 | 4.72 s |
|
||||
| Random-region control | Class probability | +0.057 | 0.077 | 0.163 | 0.023 | — |
|
||||
| Grouped occlusion | Intensity | −0.132 | 0.588 | **0.070** | −0.058 | **0.40 s** |
|
||||
| Integrated Gradients | Intensity | −0.060 | **0.632** | 0.114 | −0.036 | 4.72 s |
|
||||
| Random-region control | Intensity | −0.031 | 0.185 | 0.377 | −0.011 | — |
|
||||
|
||||
Grouped occlusion is selected for polarity: it produces a slightly larger signed class-probability drop, lower sufficiency error, higher deletion AUC, and runs about 12 times faster. For intensity, the result is a tradeoff. Integrated Gradients causes a larger prediction change when its top evidence is removed; grouped occlusion better preserves the prediction when only its top regions remain. The displayed intensity regions use grouped occlusion, with Integrated Gradients retained as a directional cross-check. The negative signed intensity changes mean that removing the selected regions raises the predicted score on average; intensity evidence is bidirectional.
|
||||
|
||||
Under standardized input noise with σ=0.02, the top-region rank Spearman correlations were 0.9993–0.9999 and top-30% Jaccard overlap was 0.987–0.998 on 120 balanced validation clips. This shows stability to that small perturbation, not stability across retrained models or a different dataset.
|
||||
|
||||
## Mapping Attachment 4 evidence to video time
|
||||
|
||||
All 20 Attachment 4 MP4 files were found, and all 20 BERT token sequences matched the supplied feature token IDs. Q1's hard CTC Viterbi word-time procedure aligned at least one word in every clip; 19 clips had full transcript word coverage, with mean word coverage 99.3%. Text wordpieces and corresponding aligned audio/vision slots inherit the transcript word interval, allowing a selected five-slot region to be shown in clip seconds.
|
||||
|
||||
CTC times are weak alignment references, not human event labels. One transcript has partial coverage. The CTC path score is uncalibrated, so it is recorded for review and is not presented as a probability or ground truth. Attachment 4's pkl files themselves do not contain Q1 `time_bounds_s`; the script computes word times from the supplied video audio and transcript.
|
||||
|
||||
Example timeline for clip 01:
|
||||
|
||||

|
||||
|
||||
## Artifacts
|
||||
|
||||
- [Explainer faithfulness and stability summary](outputs/explanation_selection/q3_explanation_method_summary.csv)
|
||||
- [Deletion/sufficiency curves](outputs/explanation_selection/q3_deletion_curves.csv)
|
||||
- [Rank stability under small input noise](outputs/explanation_selection/q3_explanation_stability.csv)
|
||||
- [Selected explanation methods and intensity tradeoff](outputs/explanation_selection/q3_explainer_selection.json)
|
||||
- [Attachment 4 predictions](outputs/explanation_selection/attachment4_predictions.csv)
|
||||
- [Top regions with word/time evidence](outputs/explanation_selection/attachment4_top_evidence.csv)
|
||||
- [Attachment 4 alignment coverage audit](outputs/explanation_selection/attachment4_alignment_audit.json)
|
||||
|
||||
The top-region CSV carries slot indices, transcript words, CTC-derived start/end seconds, uncalibrated alignment quality, prediction outputs, and modality-specific importance. It can be used to inspect individual samples or prepare the Chapter 4 evidence examples.
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 102 KiB |
@@ -0,0 +1,10 @@
|
||||
{
|
||||
"n_samples": 20,
|
||||
"n_video_files_found": 20,
|
||||
"n_ctc_any_words_aligned": 20,
|
||||
"n_ctc_full_word_coverage": 19,
|
||||
"mean_transcript_word_coverage": 0.9932742662282303,
|
||||
"n_bert_token_sequences_matching_pickle": 20,
|
||||
"time_mapping": "Q1 B1 CTC Viterbi hard word intervals computed from the supplied Attachment 4 video audio and transcript; subword slots inherit their transcript word interval",
|
||||
"quality_note": "CTC path score is uncalibrated. These intervals are localization references for interpretation, not human-annotated ground truth."
|
||||
}
|
||||
@@ -0,0 +1,21 @@
|
||||
sample_id,predicted_class,predicted_class_id,predicted_class_probability,predicted_intensity,transcript,video_file_exists,ctc_alignment_status,ctc_word_coverage,ctc_aligned_words,transcript_words,bert_token_ids_match_pickle
|
||||
01,Neutral,1,0.602049708366394,-0.12586602568626404,Replacing these wear components when replacing the timing belt is essential to ensuring the new belt performs to its mileage requirements,True,ok,1.0,21,21,True
|
||||
02,Positive,2,0.4676651060581207,0.24043764173984528,We want to live by each other’s happiness - not by each other’s misery.,True,partial,0.9285714285714286,13,14,True
|
||||
03,Negative,0,0.4883529841899872,-0.6055137515068054,"There's one lender at the moment which I think is just Bankwest who don't take rental income into account, they take rental yield into account.",True,ok,1.0,25,25,True
|
||||
04,Negative,0,0.5609740614891052,0.06384700536727905,"If I blow it at the team exercise, should I kiss my chances of cheering ""GO BLUE"" goodbye?] Absolutely not.",True,ok,1.0,20,20,True
|
||||
05,Positive,2,0.8418908715248108,1.0688594579696655,"Hi, my name is Chloe, video marketer for Red Wagon Marketing.",True,ok,1.0,11,11,True
|
||||
06,Positive,2,0.8087986707687378,0.8882662057876587,"If you're a fan of dancing in that sense, just like to watch people dance, see impressive dance moves then you might want to check out this movie solely for that",True,ok,1.0,31,31,True
|
||||
07,Positive,2,0.7762729525566101,0.8499683141708374,"As Linn’s associate editor Michael Baadke reports in our November 28 issue, attendees “enthusiastically discussed the topics of growing the hobby, the future of stamp shows, and dealers and philatelic partnerships, along with ways the leading organizations involved in the stamp hobby can work together to make it succeed and grow",True,ok,0.9803921568627451,50,51,True
|
||||
08,Positive,2,0.8577593564987183,1.020017147064209,"That brings us to tonight, the Universal Design Grand Challenge",True,ok,1.0,10,10,True
|
||||
09,Negative,0,0.9559774994850159,-1.5653929710388184,"(uhh) I did not like this movie at all, I would not recommend it",True,ok,1.0,14,14,True
|
||||
10,Negative,0,0.8787182569503784,-0.9965534806251526,(umm) And you know I really do like to see fluffy chick flicks sometimes so I'm not against that but this one was pretty terrible,True,ok,1.0,25,25,True
|
||||
11,Negative,0,0.6878972053527832,-0.6313911080360413,I would be ashamed to have made this film if I was a director,True,ok,1.0,14,14,True
|
||||
12,Negative,0,0.7832551002502441,-1.2948027849197388,"Or worse, an individual previously had good credit, but usually by no fault of their own, or perhaps by fault of their own, they have let their credit sag, and credit scores is very low.",True,ok,1.0,35,35,True
|
||||
13,Neutral,1,0.4873892366886139,0.27714914083480835,"For example, I could take a set of data and from that data, I can find a relationship between any two of the given factors or more.",True,ok,1.0,27,27,True
|
||||
14,Positive,2,0.6718549728393555,0.6427467465400696,He is the co-founder of Rossen and Vettese Limited and the former Executive Director of Uniform Final Examination (UFE) courses at Toronto's York University.,True,ok,1.0,24,24,True
|
||||
15,Positive,2,0.9403521418571472,1.3164536952972412,"However, despite their poverty, the family prioritize education because they believed in its power to transform lives",True,ok,1.0,17,17,True
|
||||
16,Negative,0,0.9597424864768982,-1.8123797178268433,"It's a terrible, this is a terrible movie",True,ok,1.0,8,8,True
|
||||
17,Positive,2,0.9568064212799072,1.3108301162719727,"Applying these four design concepts to your presentations is simple, easy and will make people think you turned into a design guru.",True,ok,1.0,22,22,True
|
||||
18,Neutral,1,0.4668891727924347,-0.5185524225234985,"-And in Denmark - the first Baltic Cod fishery has been MSC – certified -Meanwhile, the Faeroese Mackerel Fishery has been denied MSC certification based on the fact that the fishery has failed to reach an agreement on mackerel quotas with Norway and the European Union.",True,ok,0.9565217391304348,44,46,True
|
||||
19,Negative,0,0.5051047205924988,-0.09867963194847107,"People are surprisingly forgiving brands when they own up to mistakes, and unfortunately some haters out there love to point fingers and jump all over imperfections, but for the most part, people understand",True,ok,1.0,33,33,True
|
||||
20,Positive,2,0.9234767556190491,1.2413525581359863,"And of course, click in the link of the description of this video for more, and we'll have more live updates and a stock market video (wrap-up) at the end of the day today.",True,ok,1.0,34,34,True
|
||||
|
@@ -0,0 +1,178 @@
|
||||
sample_id,modality,block_index,slot_start_index,slot_end_index_exclusive,slot_indices,tokens_or_wordpieces,matched_words,word_indices,time_start_s,time_end_s,ctc_quality_uncalibrated_mean,ctc_words_covered,ctc_word_coverage_clip,class_importance,intensity_importance,class_explainer,intensity_explainer,ctc_alignment_status
|
||||
01,text,1,5,10,"5,6,7,8,9",when replacing the timing belt,when replacing the timing belt,"4,5,6,7,8",1.9625,3.2225000000000006,1.2557723595913322e-13,5,1.0,0.0734131932258606,0.08602334558963776,grouped_occlusion,grouped_occlusion,ok
|
||||
01,text,0,0,5,"1,2,3,4",replacing these wear components,Replacing these wear components,"0,1,2,3",0.2225,1.9425,1.0916685621897051e-13,4,1.0,0.053081393241882324,0.12229763716459274,grouped_occlusion,grouped_occlusion,ok
|
||||
01,text,3,15,20,"15,16,17,18,19",new belt performs to its,new belt performs to its,"14,15,16,17,18",5.202500000000001,7.1625000000000005,9.867163331619365e-14,5,1.0,0.04383492469787598,0.09824259579181671,grouped_occlusion,grouped_occlusion,ok
|
||||
01,audio,3,15,20,"15,16,17,18,19",new belt performs to its,new belt performs to its,"14,15,16,17,18",5.202500000000001,7.1625000000000005,9.867163331619365e-14,5,1.0,0.0138014554977417,0.061192527413368225,grouped_occlusion,grouped_occlusion,ok
|
||||
01,audio,1,5,10,"5,6,7,8,9",when replacing the timing belt,when replacing the timing belt,"4,5,6,7,8",1.9625,3.2225000000000006,1.2557723595913322e-13,5,1.0,0.013497352600097656,0.025103554129600525,grouped_occlusion,grouped_occlusion,ok
|
||||
01,audio,2,10,15,"10,11,12,13,14",is essential to ensuring the,is essential to ensuring the,"9,10,11,12,13",3.6025000000000005,5.1825,6.037600744192027e-14,5,1.0,0.006180107593536377,0.051678575575351715,grouped_occlusion,grouped_occlusion,ok
|
||||
01,vision,2,10,15,"10,11,12,13,14",is essential to ensuring the,is essential to ensuring the,"9,10,11,12,13",3.6025000000000005,5.1825,6.037600744192027e-14,5,1.0,0.006771266460418701,0.02889835834503174,grouped_occlusion,grouped_occlusion,ok
|
||||
01,vision,0,0,5,"1,2,3,4",replacing these wear components,Replacing these wear components,"0,1,2,3",0.2225,1.9425,1.0916685621897051e-13,4,1.0,0.003824293613433838,0.00939151644706726,grouped_occlusion,grouped_occlusion,ok
|
||||
01,vision,3,15,20,"15,16,17,18,19",new belt performs to its,new belt performs to its,"14,15,16,17,18",5.202500000000001,7.1625000000000005,9.867163331619365e-14,5,1.0,0.002546370029449463,0.02828623354434967,grouped_occlusion,grouped_occlusion,ok
|
||||
02,text,1,5,10,"5,6,7,8,9",by each other ’ s,by each other’s,"4,5,6",0.9824999999999999,1.4625,2.631730448927397e-13,3,0.9285714285714286,0.14046677947044373,0.2743243873119354,grouped_occlusion,grouped_occlusion,partial
|
||||
02,text,2,10,15,"10,12,13,14",happiness not by each,happiness not by each,"7,9,10,11",1.5025,2.4025000000000003,1.0520036340214618e-13,4,0.9285714285714286,0.12429457902908325,0.22381268441677094,grouped_occlusion,grouped_occlusion,partial
|
||||
02,text,0,0,5,"1,2,3,4",we want to live,We want to live,"0,1,2,3",0.2425,0.9425,1.6290645074933628e-13,4,0.9285714285714286,0.05786612629890442,0.0930139571428299,grouped_occlusion,grouped_occlusion,partial
|
||||
02,audio,1,5,10,"5,6,7,8,9",by each other ’ s,by each other’s,"4,5,6",0.9824999999999999,1.4625,2.631730448927397e-13,3,0.9285714285714286,0.02898383140563965,0.03369395434856415,grouped_occlusion,grouped_occlusion,partial
|
||||
02,audio,0,0,5,"1,2,3,4",we want to live,We want to live,"0,1,2,3",0.2425,0.9425,1.6290645074933628e-13,4,0.9285714285714286,0.017408668994903564,0.032251402735710144,grouped_occlusion,grouped_occlusion,partial
|
||||
02,audio,2,10,15,"10,12,13,14",happiness not by each,happiness not by each,"7,9,10,11",1.5025,2.4025000000000003,1.0520036340214618e-13,4,0.9285714285714286,0.0049620866775512695,0.005869343876838684,grouped_occlusion,grouped_occlusion,partial
|
||||
02,vision,2,10,15,"10,12,13,14",happiness not by each,happiness not by each,"7,9,10,11",1.5025,2.4025000000000003,1.0520036340214618e-13,4,0.9285714285714286,0.06942322850227356,0.14820988476276398,grouped_occlusion,grouped_occlusion,partial
|
||||
02,vision,1,5,10,"5,6,7,8,9",by each other ’ s,by each other’s,"4,5,6",0.9824999999999999,1.4625,2.631730448927397e-13,3,0.9285714285714286,0.06549379229545593,0.16973082721233368,grouped_occlusion,grouped_occlusion,partial
|
||||
02,vision,3,15,20,"15,16,17,18,19",other ’ s misery .,other’s misery.,"12,13",2.4825000000000004,3.3025000000000007,7.766181684510053e-13,2,0.9285714285714286,0.04377242922782898,0.09742923080921173,grouped_occlusion,grouped_occlusion,partial
|
||||
03,text,4,20,25,"20,21,22,23,24",t take rental income into,don't take rental income into,"13,14,15,16,17",4.522500000000001,6.1225000000000005,1.960631140554847e-11,5,1.0,0.0903124213218689,0.2127276062965393,grouped_occlusion,grouped_occlusion,ok
|
||||
03,text,5,25,30,"25,26,27,28,29","account , they take rental","account, they take rental","18,19,20,21",6.242500000000001,8.3225,6.012753101977051e-13,4,1.0,0.08981853723526001,0.18761926889419556,grouped_occlusion,grouped_occlusion,ok
|
||||
03,text,0,0,5,"1,2,3,4",there ' s one,There's one,"0,1",0.0425,0.9824999999999999,4.302478284875174e-12,2,1.0,0.049811989068984985,0.13886994123458862,grouped_occlusion,grouped_occlusion,ok
|
||||
03,audio,4,20,25,"20,21,22,23,24",t take rental income into,don't take rental income into,"13,14,15,16,17",4.522500000000001,6.1225000000000005,1.960631140554847e-11,5,1.0,0.02670571208000183,0.05279737710952759,grouped_occlusion,grouped_occlusion,ok
|
||||
03,audio,2,10,15,"10,11,12,13,14",which i think is just,which I think is just,"6,7,8,9,10",2.1825000000000006,3.2225000000000006,1.63512502285273e-11,5,1.0,0.0200861394405365,0.05790430307388306,grouped_occlusion,grouped_occlusion,ok
|
||||
03,audio,5,25,30,"25,26,27,28,29","account , they take rental","account, they take rental","18,19,20,21",6.242500000000001,8.3225,6.012753101977051e-13,4,1.0,0.017825692892074585,0.05010336637496948,grouped_occlusion,grouped_occlusion,ok
|
||||
03,vision,4,20,25,"20,21,22,23,24",t take rental income into,don't take rental income into,"13,14,15,16,17",4.522500000000001,6.1225000000000005,1.960631140554847e-11,5,1.0,0.013105422258377075,0.026691555976867676,grouped_occlusion,grouped_occlusion,ok
|
||||
03,vision,5,25,30,"25,26,27,28,29","account , they take rental","account, they take rental","18,19,20,21",6.242500000000001,8.3225,6.012753101977051e-13,4,1.0,0.013015061616897583,0.018819868564605713,grouped_occlusion,grouped_occlusion,ok
|
||||
03,vision,3,15,20,"15,16,17,18,19",bank ##west who don ',Bankwest who don't,"11,12,13",3.3025000000000007,4.702500000000001,3.167890863874076e-11,3,1.0,0.01059773564338684,0.01818716526031494,grouped_occlusion,grouped_occlusion,ok
|
||||
04,text,2,10,15,"10,11,12,13,14",should i kiss my chances,should I kiss my chances,"8,9,10,11,12",2.9025000000000003,4.062500000000001,1.5775802776191961e-10,5,1.0,0.21443644165992737,0.4394690990447998,grouped_occlusion,grouped_occlusion,ok
|
||||
04,text,0,0,5,"1,2,3,4",if i blow it,If I blow it,"0,1,2,3",0.0025000000000000005,1.4025,7.158484560430912e-12,4,1.0,0.20325228571891785,0.4211195707321167,grouped_occlusion,grouped_occlusion,ok
|
||||
04,text,4,20,25,"20,21,22,23,24",""" goodbye ? ] absolutely","BLUE"" goodbye?] Absolutely","16,17,18",5.102500000000001,6.7225,2.2125524584117753e-13,3,1.0,0.12389594316482544,0.23821038007736206,grouped_occlusion,grouped_occlusion,ok
|
||||
04,audio,3,15,20,"15,16,17,18,19","of cheering "" go blue","of cheering ""GO BLUE""","13,14,15,16",4.2225,5.5025,1.1504596424116164e-11,4,1.0,0.061087846755981445,0.1968434453010559,grouped_occlusion,grouped_occlusion,ok
|
||||
04,audio,2,10,15,"10,11,12,13,14",should i kiss my chances,should I kiss my chances,"8,9,10,11,12",2.9025000000000003,4.062500000000001,1.5775802776191961e-10,5,1.0,0.05859661102294922,0.1886153370141983,grouped_occlusion,grouped_occlusion,ok
|
||||
04,audio,4,20,25,"20,21,22,23,24",""" goodbye ? ] absolutely","BLUE"" goodbye?] Absolutely","16,17,18",5.102500000000001,6.7225,2.2125524584117753e-13,3,1.0,0.042976558208465576,0.1398433893918991,grouped_occlusion,grouped_occlusion,ok
|
||||
04,vision,2,10,15,"10,11,12,13,14",should i kiss my chances,should I kiss my chances,"8,9,10,11,12",2.9025000000000003,4.062500000000001,1.5775802776191961e-10,5,1.0,0.05346882343292236,0.12295112013816833,grouped_occlusion,grouped_occlusion,ok
|
||||
04,vision,3,15,20,"15,16,17,18,19","of cheering "" go blue","of cheering ""GO BLUE""","13,14,15,16",4.2225,5.5025,1.1504596424116164e-11,4,1.0,0.05310636758804321,0.10312038660049438,grouped_occlusion,grouped_occlusion,ok
|
||||
04,vision,1,5,10,"5,6,7,8,9","at the team exercise ,","at the team exercise,","4,5,6,7",1.5625,2.8025000000000007,9.294841968446489e-11,4,1.0,0.04525059461593628,0.11550295352935791,grouped_occlusion,grouped_occlusion,ok
|
||||
05,text,2,10,15,"10,11,12,13,14",##er for red wagon marketing,marketer for Red Wagon Marketing.,"6,7,8,9,10",3.5825000000000005,5.4625,4.2536280223424595e-13,5,1.0,0.030606567859649658,0.023810386657714844,grouped_occlusion,grouped_occlusion,ok
|
||||
05,text,0,0,5,"1,2,3,4","hi , my name","Hi, my name","0,1,2",1.2625,2.2825000000000006,9.345394556125576e-12,3,1.0,0.029161453247070312,0.012323379516601562,grouped_occlusion,grouped_occlusion,ok
|
||||
05,text,1,5,10,"5,6,7,8,9","is chloe , video market","is Chloe, video marketer","3,4,5,6",2.3225000000000002,4.022500000000001,8.972688035370665e-12,4,1.0,0.015018045902252197,0.057985544204711914,grouped_occlusion,grouped_occlusion,ok
|
||||
05,audio,1,5,10,"5,6,7,8,9","is chloe , video market","is Chloe, video marketer","3,4,5,6",2.3225000000000002,4.022500000000001,8.972688035370665e-12,4,1.0,0.022373735904693604,0.12835413217544556,grouped_occlusion,grouped_occlusion,ok
|
||||
05,audio,0,0,5,"1,2,3,4","hi , my name","Hi, my name","0,1,2",1.2625,2.2825000000000006,9.345394556125576e-12,3,1.0,0.006277620792388916,0.018590211868286133,grouped_occlusion,grouped_occlusion,ok
|
||||
05,audio,3,15,20,15,.,Marketing.,10,4.982500000000001,5.4625,1.7565947322020906e-13,1,1.0,0.003395378589630127,0.015216469764709473,grouped_occlusion,grouped_occlusion,ok
|
||||
05,vision,1,5,10,"5,6,7,8,9","is chloe , video market","is Chloe, video marketer","3,4,5,6",2.3225000000000002,4.022500000000001,8.972688035370665e-12,4,1.0,0.03512507677078247,0.09098595380783081,grouped_occlusion,grouped_occlusion,ok
|
||||
05,vision,0,0,5,"1,2,3,4","hi , my name","Hi, my name","0,1,2",1.2625,2.2825000000000006,9.345394556125576e-12,3,1.0,0.02607184648513794,0.10822361707687378,grouped_occlusion,grouped_occlusion,ok
|
||||
05,vision,3,15,20,15,.,Marketing.,10,4.982500000000001,5.4625,1.7565947322020906e-13,1,1.0,0.007893681526184082,0.017847895622253418,grouped_occlusion,grouped_occlusion,ok
|
||||
06,text,3,15,20,"15,16,17,18,19","to watch people dance ,","to watch people dance,","11,12,13,14",2.5225000000000004,3.5825000000000005,9.533516620804992e-13,4,1.0,0.0643836259841919,0.11163991689682007,grouped_occlusion,grouped_occlusion,ok
|
||||
06,text,4,20,25,"20,21,22,23,24",see impressive dance moves then,see impressive dance moves then,"15,16,17,18,19",3.7225000000000006,5.442500000000001,1.003787440114341e-12,5,1.0,0.05716830492019653,0.10976755619049072,grouped_occlusion,grouped_occlusion,ok
|
||||
06,text,2,10,15,"10,11,12,13,14","that sense , just like","that sense, just like","7,8,9,10",1.6625,2.5025000000000004,6.328522515592245e-13,4,1.0,0.05548006296157837,0.1172025203704834,grouped_occlusion,grouped_occlusion,ok
|
||||
06,audio,6,30,35,"30,31,32,33,34",out this movie solely for,out this movie solely for,"25,26,27,28,29",6.482500000000001,7.562500000000001,5.6144654493064985e-12,5,1.0,0.00836336612701416,0.047490835189819336,grouped_occlusion,grouped_occlusion,ok
|
||||
06,audio,2,10,15,"10,11,12,13,14","that sense , just like","that sense, just like","7,8,9,10",1.6625,2.5025000000000004,6.328522515592245e-13,4,1.0,0.006695687770843506,0.0182039737701416,grouped_occlusion,grouped_occlusion,ok
|
||||
06,audio,3,15,20,"15,16,17,18,19","to watch people dance ,","to watch people dance,","11,12,13,14",2.5225000000000004,3.5825000000000005,9.533516620804992e-13,4,1.0,0.005690395832061768,0.0055931806564331055,grouped_occlusion,grouped_occlusion,ok
|
||||
06,vision,2,10,15,"10,11,12,13,14","that sense , just like","that sense, just like","7,8,9,10",1.6625,2.5025000000000004,6.328522515592245e-13,4,1.0,0.024024665355682373,0.06122779846191406,grouped_occlusion,grouped_occlusion,ok
|
||||
06,vision,3,15,20,"15,16,17,18,19","to watch people dance ,","to watch people dance,","11,12,13,14",2.5225000000000004,3.5825000000000005,9.533516620804992e-13,4,1.0,0.022018134593963623,0.061392247676849365,grouped_occlusion,grouped_occlusion,ok
|
||||
06,vision,4,20,25,"20,21,22,23,24",see impressive dance moves then,see impressive dance moves then,"15,16,17,18,19",3.7225000000000006,5.442500000000001,1.003787440114341e-12,5,1.0,0.02075207233428955,0.05445140600204468,grouped_occlusion,grouped_occlusion,ok
|
||||
07,text,4,20,25,"20,21,22,23,24",“ enthusiastically discussed the topics,“enthusiastically discussed the topics,"13,14,15,16",4.362500000000001,7.5025,5.132190603623443e-13,4,0.9803921568627451,0.05012655258178711,0.07725489139556885,grouped_occlusion,grouped_occlusion,ok
|
||||
07,text,5,25,30,"25,26,27,28,29","of growing the hobby ,","of growing the hobby,","17,18,19,20",7.562500000000001,9.1225,6.009222892922947e-13,4,0.9803921568627451,0.03437221050262451,0.05201399326324463,grouped_occlusion,grouped_occlusion,ok
|
||||
07,text,9,45,50,"45,46,47,48",with ways the leading,with ways the leading,"32,33,34,35",14.3225,15.3825,4.788019704647536e-13,4,0.9803921568627451,0.03220874071121216,0.0367276668548584,grouped_occlusion,grouped_occlusion,ok
|
||||
07,audio,3,15,20,"15,17,18,19","november issue , attendees","November issue, attendees","9,11,12",3.1025000000000005,4.3425,6.891936173584844e-14,3,0.9803921568627451,0.009746789932250977,0.013783574104309082,grouped_occlusion,grouped_occlusion,ok
|
||||
07,audio,1,5,10,"5,6,7,8,9",s associate editor michael ba,Linn’s associate editor Michael Baadke,"1,2,3,4,5",0.7224999999999999,2.4025000000000003,6.973944706871624e-12,5,0.9803921568627451,0.008990466594696045,0.007673025131225586,grouped_occlusion,grouped_occlusion,ok
|
||||
07,audio,2,10,15,"10,11,12,13,14",##ad ##ke reports in our,Baadke reports in our,"5,6,7,8",2.0825000000000005,3.0825000000000005,7.431383019894585e-12,4,0.9803921568627451,0.0072026848793029785,0.008348703384399414,grouped_occlusion,grouped_occlusion,ok
|
||||
07,vision,7,35,40,"35,36,37,38,39",", and dealers and phil","shows, and dealers and philatelic","25,26,27,28,29",10.522499999999999,12.8825,1.0360054611198636e-12,5,0.9803921568627451,0.004808366298675537,0.029618024826049805,grouped_occlusion,grouped_occlusion,ok
|
||||
07,vision,3,15,20,"15,17,18,19","november issue , attendees","November issue, attendees","9,11,12",3.1025000000000005,4.3425,6.891936173584844e-14,3,0.9803921568627451,0.0032321810722351074,0.0074433088302612305,grouped_occlusion,grouped_occlusion,ok
|
||||
07,vision,2,10,15,"10,11,12,13,14",##ad ##ke reports in our,Baadke reports in our,"5,6,7,8",2.0825000000000005,3.0825000000000005,7.431383019894585e-12,4,0.9803921568627451,0.0024552345275878906,0.008127868175506592,grouped_occlusion,grouped_occlusion,ok
|
||||
08,text,0,0,5,"1,2,3,4",that brings us to,That brings us to,"0,1,2,3",0.4825,1.3425,4.985850672359178e-13,4,1.0,0.11082303524017334,0.11384594440460205,grouped_occlusion,grouped_occlusion,ok
|
||||
08,text,1,5,10,"5,6,7,8,9","tonight , the universal design","tonight, the Universal Design","4,5,6,7",1.4825,3.1225000000000005,1.5200054500503443e-13,4,1.0,0.1038968563079834,0.21401238441467285,grouped_occlusion,grouped_occlusion,ok
|
||||
08,text,2,10,15,"10,11",grand challenge,Grand Challenge,"8,9",3.3825000000000003,4.322500000000001,1.194004156129736e-14,2,1.0,0.04137396812438965,0.019734859466552734,grouped_occlusion,grouped_occlusion,ok
|
||||
08,audio,2,10,15,"10,11",grand challenge,Grand Challenge,"8,9",3.3825000000000003,4.322500000000001,1.194004156129736e-14,2,1.0,0.0033051371574401855,0.022482693195343018,grouped_occlusion,grouped_occlusion,ok
|
||||
08,audio,1,5,10,"5,6,7,8,9","tonight , the universal design","tonight, the Universal Design","4,5,6,7",1.4825,3.1225000000000005,1.5200054500503443e-13,4,1.0,0.0019735097885131836,0.06727790832519531,grouped_occlusion,grouped_occlusion,ok
|
||||
08,audio,0,0,5,"1,2,3,4",that brings us to,That brings us to,"0,1,2,3",0.4825,1.3425,4.985850672359178e-13,4,1.0,0.0007883310317993164,0.06953203678131104,grouped_occlusion,grouped_occlusion,ok
|
||||
08,vision,0,0,5,"1,2,3,4",that brings us to,That brings us to,"0,1,2,3",0.4825,1.3425,4.985850672359178e-13,4,1.0,0.004726111888885498,0.004060029983520508,grouped_occlusion,grouped_occlusion,ok
|
||||
08,vision,1,5,10,"5,6,7,8,9","tonight , the universal design","tonight, the Universal Design","4,5,6,7",1.4825,3.1225000000000005,1.5200054500503443e-13,4,1.0,0.004242956638336182,0.008825302124023438,grouped_occlusion,grouped_occlusion,ok
|
||||
08,vision,2,10,15,"10,11",grand challenge,Grand Challenge,"8,9",3.3825000000000003,4.322500000000001,1.194004156129736e-14,2,1.0,0.002269923686981201,0.007381081581115723,grouped_occlusion,grouped_occlusion,ok
|
||||
09,text,1,5,10,"5,6,7,8,9",i did not like this,I did not like this,"1,2,3,4,5",1.1824999999999999,2.2025000000000006,1.8337023088832634e-11,5,1.0,0.05766040086746216,0.3939073085784912,grouped_occlusion,grouped_occlusion,ok
|
||||
09,text,2,10,15,"10,11,12,13,14","movie at all , i","movie at all, I","6,7,8,9",2.2825000000000006,3.1025000000000005,1.7216472685121577e-10,4,1.0,0.05135905742645264,0.33904457092285156,grouped_occlusion,grouped_occlusion,ok
|
||||
09,text,0,0,5,"1,2,3,4",( uh ##h ),(uhh),0,0.7825,1.1425,6.173012658529708e-13,1,1.0,0.04096817970275879,0.3692760467529297,grouped_occlusion,grouped_occlusion,ok
|
||||
09,audio,0,0,5,"1,2,3,4",( uh ##h ),(uhh),0,0.7825,1.1425,6.173012658529708e-13,1,1.0,0.004597127437591553,0.05938518047332764,grouped_occlusion,grouped_occlusion,ok
|
||||
09,audio,1,5,10,"5,6,7,8,9",i did not like this,I did not like this,"1,2,3,4,5",1.1824999999999999,2.2025000000000006,1.8337023088832634e-11,5,1.0,0.0028375983238220215,0.0619351863861084,grouped_occlusion,grouped_occlusion,ok
|
||||
09,audio,3,15,20,"15,16,17,18",would not recommend it,would not recommend it,"10,11,12,13",3.2025000000000006,4.942500000000001,2.9955450915683833e-12,4,1.0,0.0018830299377441406,0.029448747634887695,grouped_occlusion,grouped_occlusion,ok
|
||||
09,vision,2,10,15,"10,11,12,13,14","movie at all , i","movie at all, I","6,7,8,9",2.2825000000000006,3.1025000000000005,1.7216472685121577e-10,4,1.0,0.0035586953163146973,0.010941743850708008,grouped_occlusion,grouped_occlusion,ok
|
||||
09,vision,3,15,20,"15,16,17,18",would not recommend it,would not recommend it,"10,11,12,13",3.2025000000000006,4.942500000000001,2.9955450915683833e-12,4,1.0,0.002703547477722168,0.007805347442626953,grouped_occlusion,grouped_occlusion,ok
|
||||
09,vision,0,0,5,"1,2,3,4",( uh ##h ),(uhh),0,0.7825,1.1425,6.173012658529708e-13,1,1.0,0.001896202564239502,0.0043097734451293945,grouped_occlusion,grouped_occlusion,ok
|
||||
10,text,5,25,30,"25,26,27,28,29",but this one was pretty,but this one was pretty,"19,20,21,22,23",6.742500000000001,9.1625,3.4013181527343934e-12,5,1.0,0.10162562131881714,0.3853045701980591,grouped_occlusion,grouped_occlusion,ok
|
||||
10,text,2,10,15,"10,11,12,13,14",like to see fluffy chick,like to see fluffy chick,"7,8,9,10,11",2.4825000000000004,3.8425000000000002,5.712801915197975e-12,5,1.0,0.07766342163085938,0.3524249792098999,grouped_occlusion,grouped_occlusion,ok
|
||||
10,text,3,15,20,"15,16,17,18,19",flick ##s sometimes so i,flicks sometimes so I'm,"12,13,14,15",3.9025000000000003,5.242500000000001,9.396792278478092e-09,4,1.0,0.07552039623260498,0.3311324715614319,grouped_occlusion,grouped_occlusion,ok
|
||||
10,audio,5,25,30,"25,26,27,28,29",but this one was pretty,but this one was pretty,"19,20,21,22,23",6.742500000000001,9.1625,3.4013181527343934e-12,5,1.0,0.010288834571838379,0.08406132459640503,grouped_occlusion,grouped_occlusion,ok
|
||||
10,audio,1,5,10,"5,6,7,8,9",you know i really do,you know I really do,"2,3,4,5,6",1.3425,2.4225000000000003,9.51819295384607e-13,5,1.0,0.008134961128234863,0.08679646253585815,grouped_occlusion,grouped_occlusion,ok
|
||||
10,audio,2,10,15,"10,11,12,13,14",like to see fluffy chick,like to see fluffy chick,"7,8,9,10,11",2.4825000000000004,3.8425000000000002,5.712801915197975e-12,5,1.0,0.007106482982635498,0.06629914045333862,grouped_occlusion,grouped_occlusion,ok
|
||||
10,vision,3,15,20,"15,16,17,18,19",flick ##s sometimes so i,flicks sometimes so I'm,"12,13,14,15",3.9025000000000003,5.242500000000001,9.396792278478092e-09,4,1.0,0.005765736103057861,0.02699226140975952,grouped_occlusion,grouped_occlusion,ok
|
||||
10,vision,5,25,30,"25,26,27,28,29",but this one was pretty,but this one was pretty,"19,20,21,22,23",6.742500000000001,9.1625,3.4013181527343934e-12,5,1.0,0.005424618721008301,0.004094421863555908,grouped_occlusion,grouped_occlusion,ok
|
||||
10,vision,4,20,25,"20,21,22,23,24",' m not against that,I'm not against that,"15,16,17,18",5.1825,6.6225000000000005,9.396402256608505e-09,4,1.0,0.0023061037063598633,0.0014348030090332031,grouped_occlusion,grouped_occlusion,ok
|
||||
11,text,0,0,5,"1,2,3,4",i would be ashamed,I would be ashamed,"0,1,2,3",0.5225,1.3825,2.4343081894167942e-11,4,1.0,0.17121726274490356,0.3522495925426483,grouped_occlusion,grouped_occlusion,ok
|
||||
11,text,1,5,10,"5,6,7,8,9",to have made this film,to have made this film,"4,5,6,7,8",3.2225000000000006,4.202500000000001,1.8143959640334076e-12,5,1.0,0.050980210304260254,0.09193223714828491,grouped_occlusion,grouped_occlusion,ok
|
||||
11,text,2,10,15,"10,11,12,13,14",if i was a director,if I was a director,"9,10,11,12,13",4.442500000000001,5.702500000000001,1.195583991709281e-10,5,1.0,0.03379368782043457,0.03823119401931763,grouped_occlusion,grouped_occlusion,ok
|
||||
11,audio,2,10,15,"10,11,12,13,14",if i was a director,if I was a director,"9,10,11,12,13",4.442500000000001,5.702500000000001,1.195583991709281e-10,5,1.0,0.07368618249893188,0.19450706243515015,grouped_occlusion,grouped_occlusion,ok
|
||||
11,audio,1,5,10,"5,6,7,8,9",to have made this film,to have made this film,"4,5,6,7,8",3.2225000000000006,4.202500000000001,1.8143959640334076e-12,5,1.0,0.026468276977539062,0.03296065330505371,grouped_occlusion,grouped_occlusion,ok
|
||||
11,audio,0,0,5,"1,2,3,4",i would be ashamed,I would be ashamed,"0,1,2,3",0.5225,1.3825,2.4343081894167942e-11,4,1.0,0.006807446479797363,0.017702996730804443,grouped_occlusion,grouped_occlusion,ok
|
||||
11,vision,0,0,5,"1,2,3,4",i would be ashamed,I would be ashamed,"0,1,2,3",0.5225,1.3825,2.4343081894167942e-11,4,1.0,0.019326627254486084,0.04640209674835205,grouped_occlusion,grouped_occlusion,ok
|
||||
11,vision,1,5,10,"5,6,7,8,9",to have made this film,to have made this film,"4,5,6,7,8",3.2225000000000006,4.202500000000001,1.8143959640334076e-12,5,1.0,0.006760776042938232,0.037277281284332275,grouped_occlusion,grouped_occlusion,ok
|
||||
11,vision,2,10,15,"10,11,12,13,14",if i was a director,if I was a director,"9,10,11,12,13",4.442500000000001,5.702500000000001,1.195583991709281e-10,5,1.0,0.0025706887245178223,0.00287705659866333,grouped_occlusion,grouped_occlusion,ok
|
||||
12,text,6,30,35,"30,31,32,33,34",let their credit sa ##g,"let their credit sag,","25,26,27,28",9.2625,10.7425,2.0235579326793276e-13,4,1.0,0.07106262445449829,0.21757328510284424,grouped_occlusion,grouped_occlusion,ok
|
||||
12,text,3,15,20,"15,16,17,18,19","fault of their own ,","fault of their own,","12,13,14,15",5.5825000000000005,6.3825,1.8391183721102136e-13,4,1.0,0.060820698738098145,0.17711377143859863,grouped_occlusion,grouped_occlusion,ok
|
||||
12,text,4,20,25,"20,21,22,23,24",or perhaps by fault of,or perhaps by fault of,"16,17,18,19,20",6.4225,8.1025,1.301446420968425e-12,5,1.0,0.05820423364639282,0.14634323120117188,grouped_occlusion,grouped_occlusion,ok
|
||||
12,audio,5,25,30,"25,26,27,28,29","their own , they have","their own, they have","21,22,23,24",8.1425,9.2225,2.640871901066012e-13,4,1.0,0.021765828132629395,0.05425441265106201,grouped_occlusion,grouped_occlusion,ok
|
||||
12,audio,1,5,10,"5,6,7,8,9",individual previously had good credit,"individual previously had good credit,","3,4,5,6,7",1.2625,3.4425000000000003,1.981996783456316e-13,5,1.0,0.020217180252075195,0.01235973834991455,grouped_occlusion,grouped_occlusion,ok
|
||||
12,audio,2,10,15,"10,11,12,13,14",", but usually by no","credit, but usually by no","7,8,9,10,11",3.0825000000000005,5.4625,5.096125070013053e-13,5,1.0,0.01965177059173584,0.05803334712982178,grouped_occlusion,grouped_occlusion,ok
|
||||
12,vision,0,0,5,"1,2,3,4","or worse , an","Or worse, an","0,1,2",0.5824999999999999,1.2425,1.2646860654133281e-13,3,1.0,0.00710904598236084,0.015213608741760254,grouped_occlusion,grouped_occlusion,ok
|
||||
12,vision,1,5,10,"5,6,7,8,9",individual previously had good credit,"individual previously had good credit,","3,4,5,6,7",1.2625,3.4425000000000003,1.981996783456316e-13,5,1.0,0.0066879987716674805,0.025673270225524902,grouped_occlusion,grouped_occlusion,ok
|
||||
12,vision,2,10,15,"10,11,12,13,14",", but usually by no","credit, but usually by no","7,8,9,10,11",3.0825000000000005,5.4625,5.096125070013053e-13,5,1.0,0.005474686622619629,0.0271909236907959,grouped_occlusion,grouped_occlusion,ok
|
||||
13,text,1,5,10,"5,6,7,8,9",could take a set of,could take a set of,"3,4,5,6,7",0.9624999999999999,1.8225,1.1293682043937327e-10,5,1.0,0.03829997777938843,0.08469116687774658,grouped_occlusion,grouped_occlusion,ok
|
||||
13,text,0,0,5,"1,2,3,4","for example , i","For example, I","0,1,2",0.0025000000000000005,0.9025,3.8531617620664426e-12,3,1.0,0.03210568428039551,0.017864346504211426,grouped_occlusion,grouped_occlusion,ok
|
||||
13,text,4,20,25,"20,21,22,23,24",relationship between any two of,relationship between any two of,"17,18,19,20,21",5.0825000000000005,6.782500000000001,7.070703467858767e-12,5,1.0,0.01825752854347229,0.03968304395675659,grouped_occlusion,grouped_occlusion,ok
|
||||
13,audio,4,20,25,"20,21,22,23,24",relationship between any two of,relationship between any two of,"17,18,19,20,21",5.0825000000000005,6.782500000000001,7.070703467858767e-12,5,1.0,0.01745942234992981,0.03462590277194977,grouped_occlusion,grouped_occlusion,ok
|
||||
13,audio,2,10,15,"10,11,12,13,14",data and from that data,"data and from that data,","8,9,10,11,12",1.8425,4.2625,9.633590059405473e-12,5,1.0,0.01131179928779602,0.007808178663253784,grouped_occlusion,grouped_occlusion,ok
|
||||
13,audio,0,0,5,"1,2,3,4","for example , i","For example, I","0,1,2",0.0025000000000000005,0.9025,3.8531617620664426e-12,3,1.0,0.006338447332382202,0.00720486044883728,grouped_occlusion,grouped_occlusion,ok
|
||||
14,text,2,10,15,"10,11,12,13,14",and vet ##tes ##e limited,and Vettese Limited,"6,7,8",2.0025000000000004,2.8425000000000002,4.791347081170585e-13,3,1.0,0.04705315828323364,0.09162890911102295,grouped_occlusion,grouped_occlusion,ok
|
||||
14,text,4,20,25,"20,21,22,23,24",of uniform final examination (,of Uniform Final Examination (UFE),"14,15,16,17,18",4.8425,6.9225,5.581173689339982e-13,5,1.0,0.040479302406311035,0.05692708492279053,grouped_occlusion,grouped_occlusion,ok
|
||||
14,text,0,0,5,"1,2,3,4",he is the co,He is the co-founder,"0,1,2,3",0.5025,1.3225,1.8809520903830734e-13,4,1.0,0.0321732759475708,0.08200758695602417,grouped_occlusion,grouped_occlusion,ok
|
||||
14,audio,5,25,30,"25,26,27,28,29",u ##fe ) courses at,(UFE) courses at,"18,19,20",6.742500000000001,8.362499999999999,1.0878076875022951e-13,3,1.0,0.02120727300643921,0.04397076368331909,grouped_occlusion,grouped_occlusion,ok
|
||||
14,audio,2,10,15,"10,11,12,13,14",and vet ##tes ##e limited,and Vettese Limited,"6,7,8",2.0025000000000004,2.8425000000000002,4.791347081170585e-13,3,1.0,0.020826101303100586,0.029214560985565186,grouped_occlusion,grouped_occlusion,ok
|
||||
14,audio,1,5,10,"5,6,7,8,9",- founder of ross ##en,co-founder of Rossen,"3,4,5",0.7825,1.9224999999999999,6.402270448983464e-13,3,1.0,0.017303526401519775,0.04407292604446411,grouped_occlusion,grouped_occlusion,ok
|
||||
14,vision,6,30,35,"30,31,32,33,34",toronto ' s york university,Toronto's York University.,"21,22,23",8.3825,9.6625,3.617856135575942e-12,3,1.0,0.03959810733795166,0.06995487213134766,grouped_occlusion,grouped_occlusion,ok
|
||||
14,vision,5,25,30,"25,26,27,28,29",u ##fe ) courses at,(UFE) courses at,"18,19,20",6.742500000000001,8.362499999999999,1.0878076875022951e-13,3,1.0,0.03660660982131958,0.030756652355194092,grouped_occlusion,grouped_occlusion,ok
|
||||
14,vision,4,20,25,"20,21,22,23,24",of uniform final examination (,of Uniform Final Examination (UFE),"14,15,16,17,18",4.8425,6.9225,5.581173689339982e-13,5,1.0,0.03492516279220581,0.029532790184020996,grouped_occlusion,grouped_occlusion,ok
|
||||
15,text,3,15,20,"15,16,17,18,19",believed in its power to,believed in its power to,"10,11,12,13,14",4.602500000000001,6.3825,1.3666649254190005e-11,5,1.0,0.03316134214401245,0.08834660053253174,grouped_occlusion,grouped_occlusion,ok
|
||||
15,text,2,10,15,"10,11,12,13,14",##iti ##ze education because they,prioritize education because they,"6,7,8,9",2.4225000000000003,4.522500000000001,5.053813400666634e-14,4,1.0,0.022731482982635498,0.05859649181365967,grouped_occlusion,grouped_occlusion,ok
|
||||
15,text,1,5,10,"5,6,7,8,9","poverty , the family prior","poverty, the family prioritize","3,4,5,6",1.3025,2.9825000000000004,6.197984013383666e-14,4,1.0,0.013731598854064941,0.04108023643493652,grouped_occlusion,grouped_occlusion,ok
|
||||
15,audio,0,0,5,"1,2,3,4","however , despite their","However, despite their","0,1,2",0.10250000000000001,1.2825,3.8566196054618246e-14,3,1.0,0.008936703205108643,0.06280517578125,grouped_occlusion,grouped_occlusion,ok
|
||||
15,audio,1,5,10,"5,6,7,8,9","poverty , the family prior","poverty, the family prioritize","3,4,5,6",1.3025,2.9825000000000004,6.197984013383666e-14,4,1.0,0.007372438907623291,0.04567360877990723,grouped_occlusion,grouped_occlusion,ok
|
||||
15,audio,2,10,15,"10,11,12,13,14",##iti ##ze education because they,prioritize education because they,"6,7,8,9",2.4225000000000003,4.522500000000001,5.053813400666634e-14,4,1.0,0.004992425441741943,0.07340216636657715,grouped_occlusion,grouped_occlusion,ok
|
||||
15,vision,0,0,5,"1,2,3,4","however , despite their","However, despite their","0,1,2",0.10250000000000001,1.2825,3.8566196054618246e-14,3,1.0,0.006294310092926025,0.04669678211212158,grouped_occlusion,grouped_occlusion,ok
|
||||
15,vision,2,10,15,"10,11,12,13,14",##iti ##ze education because they,prioritize education because they,"6,7,8,9",2.4225000000000003,4.522500000000001,5.053813400666634e-14,4,1.0,0.003808140754699707,0.05000507831573486,grouped_occlusion,grouped_occlusion,ok
|
||||
15,vision,1,5,10,"5,6,7,8,9","poverty , the family prior","poverty, the family prioritize","3,4,5,6",1.3025,2.9825000000000004,6.197984013383666e-14,4,1.0,0.0037418007850646973,0.05190694332122803,grouped_occlusion,grouped_occlusion,ok
|
||||
16,text,1,5,10,"5,6,7,8,9","terrible , this is a","terrible, this is a","2,3,4,5",0.8025,1.9224999999999999,6.726967737989785e-13,4,1.0,0.17667824029922485,0.8273677825927734,grouped_occlusion,grouped_occlusion,ok
|
||||
16,text,0,0,5,"1,2,3,4",it ' s a,It's a,"0,1",0.4625,0.7224999999999999,5.886124124070357e-10,2,1.0,0.08441895246505737,0.6347545385360718,grouped_occlusion,grouped_occlusion,ok
|
||||
16,text,2,10,15,"10,11",terrible movie,terrible movie,"6,7",2.0625000000000004,2.8425000000000002,3.447706729088321e-13,2,1.0,0.012092411518096924,0.09678518772125244,grouped_occlusion,grouped_occlusion,ok
|
||||
16,audio,1,5,10,"5,6,7,8,9","terrible , this is a","terrible, this is a","2,3,4,5",0.8025,1.9224999999999999,6.726967737989785e-13,4,1.0,0.005522668361663818,0.07541024684906006,grouped_occlusion,grouped_occlusion,ok
|
||||
16,audio,0,0,5,"1,2,3,4",it ' s a,It's a,"0,1",0.4625,0.7224999999999999,5.886124124070357e-10,2,1.0,0.004132866859436035,0.0680011510848999,grouped_occlusion,grouped_occlusion,ok
|
||||
16,audio,2,10,15,"10,11",terrible movie,terrible movie,"6,7",2.0625000000000004,2.8425000000000002,3.447706729088321e-13,2,1.0,0.0015830397605895996,0.02998960018157959,grouped_occlusion,grouped_occlusion,ok
|
||||
16,vision,1,5,10,"5,6,7,8,9","terrible , this is a","terrible, this is a","2,3,4,5",0.8025,1.9224999999999999,6.726967737989785e-13,4,1.0,0.0030817389488220215,0.010976672172546387,grouped_occlusion,grouped_occlusion,ok
|
||||
16,vision,0,0,5,"1,2,3,4",it ' s a,It's a,"0,1",0.4625,0.7224999999999999,5.886124124070357e-10,2,1.0,0.0018818974494934082,0.008708953857421875,grouped_occlusion,grouped_occlusion,ok
|
||||
16,vision,2,10,15,"10,11",terrible movie,terrible movie,"6,7",2.0625000000000004,2.8425000000000002,3.447706729088321e-13,2,1.0,0.0007146596908569336,0.0037450790405273438,grouped_occlusion,grouped_occlusion,ok
|
||||
17,text,3,15,20,"15,16,17,18,19",make people think you turned,make people think you turned,"13,14,15,16,17",6.982500000000001,8.202499999999999,5.0308304651950566e-14,5,1.0,0.03196984529495239,0.10473096370697021,grouped_occlusion,grouped_occlusion,ok
|
||||
17,text,4,20,25,"20,21,22,23,24",into a design guru .,into a design guru.,"18,19,20,21",8.3025,9.5425,6.876211439212395e-12,4,1.0,0.025771260261535645,0.09248185157775879,grouped_occlusion,grouped_occlusion,ok
|
||||
17,text,2,10,15,"10,11,12,13,14","simple , easy and will","simple, easy and will","9,10,11,12",4.8025,6.942500000000001,8.080162655757211e-13,4,1.0,0.02379608154296875,0.15895986557006836,grouped_occlusion,grouped_occlusion,ok
|
||||
17,audio,0,0,5,"1,2,3,4",applying these four design,Applying these four design,"0,1,2,3",0.0025000000000000005,2.1225000000000005,3.907075057276878e-13,4,1.0,0.001038670539855957,0.02217411994934082,grouped_occlusion,grouped_occlusion,ok
|
||||
17,audio,1,5,10,"5,6,7,8,9",concepts to your presentations is,concepts to your presentations is,"4,5,6,7,8",2.1825000000000006,4.4225,1.3654575542430864e-13,5,1.0,0.0008327364921569824,0.029980182647705078,grouped_occlusion,grouped_occlusion,ok
|
||||
17,audio,4,20,25,"20,21,22,23,24",into a design guru .,into a design guru.,"18,19,20,21",8.3025,9.5425,6.876211439212395e-12,4,1.0,0.0007992982864379883,0.01889348030090332,grouped_occlusion,grouped_occlusion,ok
|
||||
17,vision,3,15,20,"15,16,17,18,19",make people think you turned,make people think you turned,"13,14,15,16,17",6.982500000000001,8.202499999999999,5.0308304651950566e-14,5,1.0,0.003390192985534668,0.05066335201263428,grouped_occlusion,grouped_occlusion,ok
|
||||
17,vision,2,10,15,"10,11,12,13,14","simple , easy and will","simple, easy and will","9,10,11,12",4.8025,6.942500000000001,8.080162655757211e-13,4,1.0,0.0014863014221191406,0.04471385478973389,grouped_occlusion,grouped_occlusion,ok
|
||||
17,vision,4,20,25,"20,21,22,23,24",into a design guru .,into a design guru.,"18,19,20,21",8.3025,9.5425,6.876211439212395e-12,4,1.0,0.0008729696273803711,0.02221369743347168,grouped_occlusion,grouped_occlusion,ok
|
||||
18,text,0,0,5,"1,2,3,4",- and in denmark,-And in Denmark,"0,1,2",0.5625,1.2025,7.339091369947377e-14,3,0.9565217391304348,0.040209442377090454,0.07814651727676392,grouped_occlusion,grouped_occlusion,ok
|
||||
18,text,1,5,10,"6,7,8,9",the first baltic cod,the first Baltic Cod,"4,5,6,7",1.2825,2.7225000000000006,1.9238576151557236e-12,4,0.9565217391304348,0.03645741939544678,0.07523351907730103,grouped_occlusion,grouped_occlusion,ok
|
||||
18,text,7,35,40,"35,36,37,38,39",fact that the fishery has,fact that the fishery has,"27,28,29,30,31",11.4625,12.6625,3.051134360523701e-14,5,0.9565217391304348,0.02193853259086609,0.09061205387115479,grouped_occlusion,grouped_occlusion,ok
|
||||
18,audio,1,5,10,"6,7,8,9",the first baltic cod,the first Baltic Cod,"4,5,6,7",1.2825,2.7225000000000006,1.9238576151557236e-12,4,0.9565217391304348,0.011634111404418945,0.011849582195281982,grouped_occlusion,grouped_occlusion,ok
|
||||
18,audio,3,15,20,"15,16,17,18,19","certified - meanwhile , the","certified -Meanwhile, the","13,14,15",4.522500000000001,7.242500000000001,2.0090313043947478e-13,3,0.9565217391304348,0.009550690650939941,0.00717240571975708,grouped_occlusion,grouped_occlusion,ok
|
||||
18,audio,4,20,25,"20,21,22,23,24",fae ##ro ##ese mack ##ere,Faeroese Mackerel,"16,17",7.2625,8.1625,7.552127939731886e-13,2,0.9565217391304348,0.0075453221797943115,0.0038139820098876953,grouped_occlusion,grouped_occlusion,ok
|
||||
18,vision,3,15,20,"15,16,17,18,19","certified - meanwhile , the","certified -Meanwhile, the","13,14,15",4.522500000000001,7.242500000000001,2.0090313043947478e-13,3,0.9565217391304348,0.015370607376098633,0.00964266061782837,grouped_occlusion,grouped_occlusion,ok
|
||||
18,vision,4,20,25,"20,21,22,23,24",fae ##ro ##ese mack ##ere,Faeroese Mackerel,"16,17",7.2625,8.1625,7.552127939731886e-13,2,0.9565217391304348,0.01368647813796997,0.025389909744262695,grouped_occlusion,grouped_occlusion,ok
|
||||
18,vision,7,35,40,"35,36,37,38,39",fact that the fishery has,fact that the fishery has,"27,28,29,30,31",11.4625,12.6625,3.051134360523701e-14,5,0.9565217391304348,0.012452512979507446,0.0016064047813415527,grouped_occlusion,grouped_occlusion,ok
|
||||
19,text,5,25,30,"25,26,27,28,29",and jump all over imperfect,"and jump all over imperfections,","21,22,23,24,25",8.9225,10.7025,2.4467552179386886e-13,5,1.0,0.14123356342315674,0.26041585206985474,grouped_occlusion,grouped_occlusion,ok
|
||||
19,text,2,10,15,"10,11,12,13,14","up to mistakes , and","up to mistakes, and","8,9,10,11",3.6025000000000005,5.2625,2.756182849760793e-14,4,1.0,0.1259310245513916,0.27155396342277527,grouped_occlusion,grouped_occlusion,ok
|
||||
19,text,1,5,10,"5,6,7,8,9",##giving brands when they own,forgiving brands when they own,"3,4,5,6,7",1.8025,3.4025000000000003,4.482843688564828e-15,5,1.0,0.10481253266334534,0.20342162251472473,grouped_occlusion,grouped_occlusion,ok
|
||||
19,audio,5,25,30,"25,26,27,28,29",and jump all over imperfect,"and jump all over imperfections,","21,22,23,24,25",8.9225,10.7025,2.4467552179386886e-13,5,1.0,0.0375896692276001,0.0982179045677185,grouped_occlusion,grouped_occlusion,ok
|
||||
19,audio,7,35,40,"35,36,37,38,39","most part , people understand","most part, people understand","29,30,31,32",11.8425,13.5625,6.956549052765396e-14,4,1.0,0.03040042519569397,0.054060935974121094,grouped_occlusion,grouped_occlusion,ok
|
||||
19,audio,4,20,25,"20,21,22,23,24",there love to point fingers,there love to point fingers,"16,17,18,19,20",6.942500000000001,8.862499999999999,2.039331089220153e-13,5,1.0,0.021963000297546387,0.04821614921092987,grouped_occlusion,grouped_occlusion,ok
|
||||
19,vision,3,15,20,"15,16,17,18,19",unfortunately some hate ##rs out,unfortunately some haters out,"12,13,14,15",5.3425,6.9225,4.614146997521066e-14,4,1.0,0.031420767307281494,0.07558748126029968,grouped_occlusion,grouped_occlusion,ok
|
||||
19,vision,1,5,10,"5,6,7,8,9",##giving brands when they own,forgiving brands when they own,"3,4,5,6,7",1.8025,3.4025000000000003,4.482843688564828e-15,5,1.0,0.029108166694641113,0.07041062414646149,grouped_occlusion,grouped_occlusion,ok
|
||||
19,vision,5,25,30,"25,26,27,28,29",and jump all over imperfect,"and jump all over imperfections,","21,22,23,24,25",8.9225,10.7025,2.4467552179386886e-13,5,1.0,0.02759939432144165,0.028621017932891846,grouped_occlusion,grouped_occlusion,ok
|
||||
20,text,7,35,40,"35,36,37,38,39",) at the end of,(wrap-up) at the end of,"26,27,28,29,30",8.5025,10.1025,1.1233773471618156e-12,5,1.0,0.022437691688537598,0.02704167366027832,grouped_occlusion,grouped_occlusion,ok
|
||||
20,text,0,0,5,"1,2,3,4","and of course ,","And of course,","0,1,2",0.0225,1.0425,7.323313211132221e-13,3,1.0,0.01609170436859131,0.03773140907287598,grouped_occlusion,grouped_occlusion,ok
|
||||
20,text,5,25,30,"25,26,27,28,29",updates and a stock market,updates and a stock market,"20,21,22,23,24",6.6625000000000005,8.1625,6.493302127844331e-12,5,1.0,0.01604229211807251,0.03772282600402832,grouped_occlusion,grouped_occlusion,ok
|
||||
20,audio,8,40,45,"40,41,42,43",the day today .,the day today.,"31,32,33",10.1425,10.8025,5.934335502395015e-12,3,1.0,0.0015968680381774902,0.016044139862060547,grouped_occlusion,grouped_occlusion,ok
|
||||
20,audio,3,15,20,"15,16,17,18,19","for more , and we","for more, and we'll","13,14,15,16",4.3825,5.522500000000001,1.612313071618996e-11,4,1.0,0.0011889338493347168,0.01394963264465332,grouped_occlusion,grouped_occlusion,ok
|
||||
20,audio,1,5,10,"5,6,7,8,9",click in the link of,click in the link of,"3,4,5,6,7",1.1025,2.1425000000000005,2.2686710998425595e-12,5,1.0,0.0008327364921569824,0.015181779861450195,grouped_occlusion,grouped_occlusion,ok
|
||||
20,vision,2,10,15,"10,11,12,13,14",the description of this video,the description of this video,"8,9,10,11,12",2.2425000000000006,3.8225000000000007,1.1555191343633454e-11,5,1.0,0.006613016128540039,0.03737950325012207,grouped_occlusion,grouped_occlusion,ok
|
||||
20,vision,3,15,20,"15,16,17,18,19","for more , and we","for more, and we'll","13,14,15,16",4.3825,5.522500000000001,1.612313071618996e-11,4,1.0,0.004205465316772461,0.032741665840148926,grouped_occlusion,grouped_occlusion,ok
|
||||
20,vision,1,5,10,"5,6,7,8,9",click in the link of,click in the link of,"3,4,5,6,7",1.1025,2.1425000000000005,2.2686710998425595e-12,5,1.0,0.0037073493003845215,0.016587018966674805,grouped_occlusion,grouped_occlusion,ok
|
||||
|
@@ -0,0 +1,37 @@
|
||||
method,target,fraction_removed,mean_signed_drop,mean_absolute_change,n_valid
|
||||
integrated_gradients,predicted_class_probability,0.1,0.1318911910057068,0.1598358154296875,728
|
||||
integrated_gradients,predicted_class_probability,0.2,0.2255048155784607,0.2508196234703064,728
|
||||
integrated_gradients,predicted_class_probability,0.3,0.2575814723968506,0.2828781306743622,728
|
||||
integrated_gradients,predicted_class_probability,0.4,0.2559848725795746,0.28699347376823425,728
|
||||
integrated_gradients,predicted_class_probability,0.5,0.25572627782821655,0.29257047176361084,728
|
||||
integrated_gradients,predicted_class_probability,-0.3,0.039929818361997604,0.039929818361997604,728
|
||||
integrated_gradients,intensity,0.1,-0.04874260723590851,0.37119486927986145,728
|
||||
integrated_gradients,intensity,0.2,-0.03821782395243645,0.5629051327705383,728
|
||||
integrated_gradients,intensity,0.3,-0.05994963273406029,0.6319279074668884,728
|
||||
integrated_gradients,intensity,0.4,-0.13912875950336456,0.653501570224762,728
|
||||
integrated_gradients,intensity,0.5,-0.20183886587619781,0.6705618500709534,728
|
||||
integrated_gradients,intensity,-0.3,0.1144869476556778,0.1144869476556778,728
|
||||
grouped_occlusion,predicted_class_probability,0.1,0.18875761330127716,0.20476296544075012,728
|
||||
grouped_occlusion,predicted_class_probability,0.2,0.2506031095981598,0.2624903619289398,728
|
||||
grouped_occlusion,predicted_class_probability,0.3,0.26046913862228394,0.27427902817726135,728
|
||||
grouped_occlusion,predicted_class_probability,0.4,0.2578567862510681,0.27872106432914734,728
|
||||
grouped_occlusion,predicted_class_probability,0.5,0.2540586292743683,0.2836124897003174,728
|
||||
grouped_occlusion,predicted_class_probability,-0.3,0.024163084104657173,0.024163084104657173,728
|
||||
grouped_occlusion,intensity,0.1,-0.056311462074518204,0.4420826733112335,728
|
||||
grouped_occlusion,intensity,0.2,-0.08403468877077103,0.5550013184547424,728
|
||||
grouped_occlusion,intensity,0.3,-0.13223451375961304,0.5881773829460144,728
|
||||
grouped_occlusion,intensity,0.4,-0.20568883419036865,0.6224051117897034,728
|
||||
grouped_occlusion,intensity,0.5,-0.2683386206626892,0.6487718820571899,728
|
||||
grouped_occlusion,intensity,-0.3,0.0703125074505806,0.0703125074505806,728
|
||||
random,predicted_class_probability,0.1,0.017329903319478035,0.03110884316265583,728
|
||||
random,predicted_class_probability,0.2,0.0376301035284996,0.05323585495352745,728
|
||||
random,predicted_class_probability,0.3,0.057446520775556564,0.07715633511543274,728
|
||||
random,predicted_class_probability,0.4,0.07688438147306442,0.09651552885770798,728
|
||||
random,predicted_class_probability,0.5,0.0910244807600975,0.11337775737047195,728
|
||||
random,predicted_class_probability,-0.3,0.16268795728683472,0.16268795728683472,728
|
||||
random,intensity,0.1,-0.01041108462959528,0.07401501387357712,728
|
||||
random,intensity,0.2,-0.012165687046945095,0.12958525121212006,728
|
||||
random,intensity,0.3,-0.031106332316994667,0.18503554165363312,728
|
||||
random,intensity,0.4,-0.03734554722905159,0.2288973033428192,728
|
||||
random,intensity,0.5,-0.04232712835073471,0.2670633792877197,728
|
||||
random,intensity,-0.3,0.37735068798065186,0.37735068798065186,728
|
||||
|
@@ -0,0 +1,18 @@
|
||||
{
|
||||
"q2_predictor": "concat",
|
||||
"classification_explainer": "grouped_occlusion",
|
||||
"intensity_explainer_primary": "grouped_occlusion",
|
||||
"intensity_explainer_crosscheck": "integrated_gradients",
|
||||
"intensity_tradeoff": {
|
||||
"integrated_gradients_abs_change_at_30": 0.6319279074668884,
|
||||
"integrated_gradients_sufficiency_error_top_30": 0.1144869476556778,
|
||||
"grouped_occlusion_abs_change_at_30": 0.5881773829460144,
|
||||
"grouped_occlusion_sufficiency_error_top_30": 0.0703125074505806
|
||||
},
|
||||
"selection_basis": "For polarity, grouped occlusion has the larger signed target-probability drop, lower sufficiency error, higher deletion AUC, and lower runtime. For intensity, IG causes a larger deletion change but grouped occlusion has lower top-evidence sufficiency error; grouped occlusion is used for displayed segments and IG is retained as a cross-check. No combined explanation score is used.",
|
||||
"valid_samples": 728,
|
||||
"stability_samples": 120,
|
||||
"integrated_gradients_runtime_seconds": 4.718544340998051,
|
||||
"grouped_occlusion_runtime_seconds": 0.40209812500688713,
|
||||
"ctc_time_map_for_attachment4": "Q1 hard CTC Viterbi word boundaries from source video audio; not human alignment ground truth"
|
||||
}
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 138 KiB |
@@ -0,0 +1,7 @@
|
||||
method,target,comprehensiveness_signed_drop_at_30,absolute_prediction_change_at_30,sufficiency_abs_error_top_30,deletion_drop_auc_10_to_50,runtime_seconds,spearman_rank_correlation,top_30_percent_jaccard,n_samples,input_noise_sigma
|
||||
integrated_gradients,predicted_class_probability,0.2575814723968506,0.2828781306743622,0.039929818361997604,0.09328798949718475,4.718544340998051,0.9998964070157896,0.9966666666666666,120,0.02
|
||||
integrated_gradients,intensity,-0.05994963273406029,0.6319279074668884,0.1144869476556778,-0.03625869527459145,4.718544340998051,0.9999221890615111,0.9983333333333333,120,0.02
|
||||
grouped_occlusion,predicted_class_probability,0.26046913862228394,0.27427902817726135,0.024163084104657173,0.09903371557593346,0.40209812500688713,0.9993435630984816,0.99,120,0.02
|
||||
grouped_occlusion,intensity,-0.13223451375961304,0.5881773829460144,0.0703125074505806,-0.05842830780893564,0.40209812500688713,0.9994847562909268,0.9866666666666667,120,0.02
|
||||
random,predicted_class_probability,0.057446520775556564,0.07715633511543274,0.16268795728683472,0.022613819781690837,0.40209812500688713,nan,nan,,
|
||||
random,intensity,-0.031106332316994667,0.18503554165363312,0.37735068798065186,-0.010698667308315635,0.40209812500688713,nan,nan,,
|
||||
|
@@ -0,0 +1,5 @@
|
||||
method,target,spearman_rank_correlation,top_30_percent_jaccard,n_samples,input_noise_sigma
|
||||
integrated_gradients,predicted_class_probability,0.9998964070157896,0.9966666666666666,120,0.02
|
||||
grouped_occlusion,predicted_class_probability,0.9993435630984816,0.99,120,0.02
|
||||
integrated_gradients,intensity,0.9999221890615111,0.9983333333333333,120,0.02
|
||||
grouped_occlusion,intensity,0.9994847562909268,0.9866666666666667,120,0.02
|
||||
|
@@ -0,0 +1 @@
|
||||
"""Q3 interpretation and evidence-localization experiments."""
|
||||
@@ -0,0 +1,148 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import subprocess
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
from transformers import AutoModelForCTC, AutoTokenizer
|
||||
|
||||
|
||||
SAMPLE_RATE = 16_000
|
||||
MODEL_ID = "facebook/wav2vec2-base-960h"
|
||||
|
||||
|
||||
@dataclass
|
||||
class WordInterval:
|
||||
word: str
|
||||
start_s: float
|
||||
end_s: float
|
||||
quality: float
|
||||
valid: bool
|
||||
|
||||
|
||||
def decode_audio(video_path: Path) -> np.ndarray:
|
||||
result = subprocess.run(
|
||||
[
|
||||
"ffmpeg", "-v", "error", "-i", str(video_path), "-map", "0:a:0",
|
||||
"-ac", "1", "-ar", str(SAMPLE_RATE), "-f", "f32le", "pipe:1",
|
||||
],
|
||||
check=True,
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.PIPE,
|
||||
)
|
||||
waveform = np.frombuffer(result.stdout, dtype="<f4").copy()
|
||||
if not len(waveform):
|
||||
raise RuntimeError(f"no decoded audio in {video_path}")
|
||||
return waveform
|
||||
|
||||
|
||||
def _ctc_targets(words: list[str], tokenizer) -> tuple[list[int], list[list[int]]]:
|
||||
vocab = tokenizer.get_vocab()
|
||||
delimiter = int(tokenizer.convert_tokens_to_ids(tokenizer.word_delimiter_token or "|"))
|
||||
unknown = int(tokenizer.unk_token_id)
|
||||
targets: list[int] = []
|
||||
per_word: list[list[int]] = [[] for _ in words]
|
||||
for word_index, raw_word in enumerate(words):
|
||||
if word_index:
|
||||
targets.append(delimiter)
|
||||
normalized = re.sub(r"[^a-z']", "", raw_word.lower())
|
||||
for character in normalized:
|
||||
per_word[word_index].append(len(targets))
|
||||
targets.append(int(vocab.get(character, unknown)))
|
||||
return targets, per_word
|
||||
|
||||
|
||||
def _viterbi(log_probs: np.ndarray, targets: list[int], blank: int) -> np.ndarray | None:
|
||||
if not targets or log_probs.ndim != 2:
|
||||
return None
|
||||
states = np.full(2 * len(targets) + 1, blank, dtype=np.int64)
|
||||
states[1::2] = np.asarray(targets, dtype=np.int64)
|
||||
frames, count = log_probs.shape[0], len(states)
|
||||
if frames == 0 or frames < len(targets):
|
||||
return None
|
||||
previous = np.full(count, -np.inf, dtype=np.float64)
|
||||
previous[0] = float(log_probs[0, blank])
|
||||
previous[1] = float(log_probs[0, states[1]])
|
||||
back = np.zeros((frames, count), dtype=np.uint8)
|
||||
skip = np.zeros(count, dtype=bool)
|
||||
if count > 2:
|
||||
skip[2:] = (states[2:] != blank) & (states[2:] != states[:-2])
|
||||
for frame in range(1, frames):
|
||||
stay = previous
|
||||
one = np.full(count, -np.inf, dtype=np.float64)
|
||||
one[1:] = previous[:-1]
|
||||
two = np.full(count, -np.inf, dtype=np.float64)
|
||||
if skip.any():
|
||||
two[skip] = previous[np.flatnonzero(skip) - 2]
|
||||
candidates = np.stack((stay, one, two), axis=0)
|
||||
choice = candidates.argmax(axis=0).astype(np.uint8)
|
||||
previous = candidates[choice, np.arange(count)] + log_probs[frame, states]
|
||||
back[frame] = choice
|
||||
state = count - 1 if previous[-1] >= previous[-2] else count - 2
|
||||
path = np.empty(frames, dtype=np.int32)
|
||||
path[-1] = state
|
||||
for frame in range(frames - 1, 0, -1):
|
||||
state -= int(back[frame, state])
|
||||
path[frame - 1] = state
|
||||
return path
|
||||
|
||||
|
||||
def model_time_constants(model) -> tuple[float, float]:
|
||||
config = model.config
|
||||
stride = int(np.prod(config.conv_stride))
|
||||
receptive = 1
|
||||
jump = 1
|
||||
for kernel, local_stride in zip(config.conv_kernel, config.conv_stride):
|
||||
receptive += (int(kernel) - 1) * jump
|
||||
jump *= int(local_stride)
|
||||
return stride / SAMPLE_RATE, receptive / (2 * SAMPLE_RATE)
|
||||
|
||||
|
||||
def align_words(
|
||||
waveform: np.ndarray,
|
||||
words: list[str],
|
||||
tokenizer,
|
||||
model,
|
||||
device: torch.device,
|
||||
) -> list[WordInterval]:
|
||||
targets, word_targets = _ctc_targets(words, tokenizer)
|
||||
blank = int(tokenizer.pad_token_id)
|
||||
if not targets or not len(waveform):
|
||||
return [WordInterval(w, float("nan"), float("nan"), 0.0, False) for w in words]
|
||||
with torch.inference_mode():
|
||||
values = torch.as_tensor(waveform, dtype=torch.float32, device=device).unsqueeze(0)
|
||||
logits = model(input_values=values).logits[0].float()
|
||||
log_probs = torch.log_softmax(logits, dim=-1).cpu().numpy()
|
||||
path = _viterbi(log_probs, targets, blank)
|
||||
frame_step, center_s = model_time_constants(model)
|
||||
duration = len(waveform) / SAMPLE_RATE
|
||||
intervals: list[WordInterval] = []
|
||||
if path is None:
|
||||
return [WordInterval(w, float("nan"), float("nan"), 0.0, False) for w in words]
|
||||
for word, target_indices in zip(words, word_targets):
|
||||
states = np.asarray([2 * index + 1 for index in target_indices], dtype=np.int32)
|
||||
frame_indices = np.flatnonzero(np.isin(path, states)) if len(states) else np.empty(0, dtype=np.int64)
|
||||
if not len(frame_indices):
|
||||
intervals.append(WordInterval(word, float("nan"), float("nan"), 0.0, False))
|
||||
continue
|
||||
first, last = int(frame_indices[0]), int(frame_indices[-1])
|
||||
start = max(0.0, first * frame_step + center_s - frame_step / 2)
|
||||
end = min(duration, (last + 1) * frame_step + center_s - frame_step / 2)
|
||||
char_scores = []
|
||||
for target_index in target_indices:
|
||||
selected = np.flatnonzero(path == 2 * target_index + 1)
|
||||
if len(selected):
|
||||
char_scores.extend(log_probs[selected, targets[target_index]].tolist())
|
||||
quality = float(np.exp(np.mean(char_scores))) if char_scores else 0.0
|
||||
valid = end > start
|
||||
intervals.append(WordInterval(word, start, end, quality, valid))
|
||||
return intervals
|
||||
|
||||
|
||||
def load_ctc(device: torch.device):
|
||||
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
|
||||
model = AutoModelForCTC.from_pretrained(MODEL_ID).to(device).eval()
|
||||
return tokenizer, model
|
||||
@@ -0,0 +1,622 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import csv
|
||||
import json
|
||||
import pickle
|
||||
import re
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
Q2_PROJECT = Path(__file__).resolve().parents[2] / "Q2"
|
||||
sys.path.insert(0, str(Q2_PROJECT))
|
||||
|
||||
import matplotlib
|
||||
matplotlib.use("Agg")
|
||||
import matplotlib.pyplot as plt
|
||||
import numpy as np
|
||||
import torch
|
||||
from scipy.stats import spearmanr
|
||||
from transformers import AutoTokenizer
|
||||
|
||||
from .ctc_time import align_words, decode_audio, load_ctc
|
||||
from q2.data import MODALITIES, ROOT, RobustStats, Split, apply_robust_stats, load_aligned
|
||||
from q2.models import AlignedFusionModel
|
||||
from q2.train_compare import _score_arrays, _write_csv
|
||||
|
||||
|
||||
ATTACHMENT4 = ROOT / "E题数据" / "附件4-可解释专项视频样本与特征文件" / "附件4-可解释专项视频样本与特征文件" / "对齐版本"
|
||||
MODALITY_LABELS = {0: "text", 1: "audio", 2: "vision"}
|
||||
CLASS_NAMES = {0: "Negative", 1: "Neutral", 2: "Positive"}
|
||||
BLOCK = 5
|
||||
N_BLOCKS = 50 // BLOCK
|
||||
|
||||
|
||||
def _model_from_run(output: Path, device: torch.device):
|
||||
method = (output / "selected_method.txt").read_text(encoding="utf-8").split(":", 1)[1].split(".", 1)[0].strip()
|
||||
checkpoint = torch.load(output / "models" / "aligned" / method / "model_best.pt", map_location=device, weights_only=False)
|
||||
model = AlignedFusionModel(method, tuple(checkpoint["dims"])).to(device)
|
||||
model.load_state_dict(checkpoint["state_dict"])
|
||||
model.eval()
|
||||
return method, model
|
||||
|
||||
|
||||
def _selected_scores(
|
||||
model: AlignedFusionModel,
|
||||
split: Split,
|
||||
mask: np.ndarray,
|
||||
device: torch.device,
|
||||
batch: int = 128,
|
||||
target_classes: np.ndarray | None = None,
|
||||
):
|
||||
model.eval()
|
||||
all_logits, all_reg = [], []
|
||||
with torch.inference_mode():
|
||||
for start in range(0, split.n, batch):
|
||||
stop = min(start + batch, split.n)
|
||||
xs = tuple(torch.as_tensor(x[start:stop], dtype=torch.float32, device=device) for x in split.x)
|
||||
mb = torch.as_tensor(mask[start:stop], dtype=torch.bool, device=device)
|
||||
result = model(xs, mb)
|
||||
all_logits.append(result["logits"].float().cpu().numpy())
|
||||
all_reg.append(result["intensity"].float().cpu().numpy())
|
||||
logits = np.concatenate(all_logits)
|
||||
intensity = np.clip(np.concatenate(all_reg), -3.0, 3.0)
|
||||
pred_class = logits.argmax(axis=-1)
|
||||
selected_class = pred_class if target_classes is None else np.asarray(target_classes, dtype=np.int64)
|
||||
prob = torch.softmax(torch.as_tensor(logits), dim=-1).numpy()[np.arange(split.n), selected_class]
|
||||
return logits, pred_class, prob, intensity
|
||||
|
||||
|
||||
def _integrated_groups(
|
||||
model: AlignedFusionModel,
|
||||
split: Split,
|
||||
masks: np.ndarray,
|
||||
device: torch.device,
|
||||
steps: int = 16,
|
||||
batch_size: int = 48,
|
||||
) -> tuple[np.ndarray, np.ndarray]:
|
||||
"""Absolute Integrated Gradients grouped into three modalities x ten 5-slot blocks."""
|
||||
model.eval()
|
||||
class_scores = np.zeros((split.n, 3, N_BLOCKS), dtype=np.float32)
|
||||
reg_scores = np.zeros_like(class_scores)
|
||||
for start in range(0, split.n, batch_size):
|
||||
stop = min(start + batch_size, split.n)
|
||||
xb = tuple(torch.as_tensor(x[start:stop], dtype=torch.float32, device=device) for x in split.x)
|
||||
mb = torch.as_tensor(masks[start:stop], dtype=torch.bool, device=device)
|
||||
with torch.no_grad():
|
||||
base = model(xb, mb)
|
||||
target = base["logits"].argmax(dim=-1)
|
||||
grad_class = [torch.zeros_like(x) for x in xb]
|
||||
grad_reg = [torch.zeros_like(x) for x in xb]
|
||||
# cuDNN's fused GRU does not support backward while the module is in
|
||||
# eval mode; the non-fused implementation is mathematically identical.
|
||||
with torch.backends.cudnn.flags(enabled=False):
|
||||
for alpha in torch.linspace(1.0 / steps, 1.0, steps, device=device):
|
||||
inputs = tuple((x * alpha).detach().requires_grad_(True) for x in xb)
|
||||
output = model(inputs, mb)
|
||||
target_prob = torch.softmax(output["logits"], dim=-1).gather(1, target[:, None]).sum()
|
||||
gradients = torch.autograd.grad(target_prob, inputs, retain_graph=True)
|
||||
reg_gradients = torch.autograd.grad(output["intensity"].sum(), inputs)
|
||||
for modality in range(3):
|
||||
grad_class[modality] += gradients[modality].detach()
|
||||
grad_reg[modality] += reg_gradients[modality].detach()
|
||||
for modality in range(3):
|
||||
attr_class = (xb[modality] * grad_class[modality] / steps).abs().sum(dim=-1)
|
||||
attr_reg = (xb[modality] * grad_reg[modality] / steps).abs().sum(dim=-1)
|
||||
attr_class = attr_class.reshape(stop - start, N_BLOCKS, BLOCK).sum(dim=-1)
|
||||
attr_reg = attr_reg.reshape(stop - start, N_BLOCKS, BLOCK).sum(dim=-1)
|
||||
class_scores[start:stop, modality] = attr_class.float().cpu().numpy()
|
||||
reg_scores[start:stop, modality] = attr_reg.float().cpu().numpy()
|
||||
print(f"[IG] explained validation rows {start}:{stop}/{split.n}", flush=True)
|
||||
return class_scores, reg_scores
|
||||
|
||||
|
||||
def _occlusion_groups(
|
||||
model: AlignedFusionModel,
|
||||
split: Split,
|
||||
masks: np.ndarray,
|
||||
device: torch.device,
|
||||
) -> tuple[np.ndarray, np.ndarray]:
|
||||
"""Measure the prediction change when one aligned five-slot modality block is hidden."""
|
||||
full_logits, full_class, full_prob, full_reg = _selected_scores(model, split, masks, device)
|
||||
class_scores = np.zeros((split.n, 3, N_BLOCKS), dtype=np.float32)
|
||||
reg_scores = np.zeros_like(class_scores)
|
||||
for modality in range(3):
|
||||
for block in range(N_BLOCKS):
|
||||
changed = masks.copy()
|
||||
left, right = block * BLOCK, (block + 1) * BLOCK
|
||||
changed[:, left:right, modality] = False
|
||||
_, _, prob, reg = _selected_scores(model, split, changed, device, target_classes=full_class)
|
||||
class_scores[:, modality, block] = np.abs(full_prob - prob)
|
||||
reg_scores[:, modality, block] = np.abs(full_reg - reg)
|
||||
print(f"[occlusion] finished {MODALITY_LABELS[modality]}", flush=True)
|
||||
return class_scores, reg_scores
|
||||
|
||||
|
||||
def _rank_delete_masks(base: np.ndarray, scores: np.ndarray, fraction: float, keep: bool = False) -> np.ndarray:
|
||||
n, steps, modalities = base.shape
|
||||
count = max(1, int(round(fraction * 3 * N_BLOCKS)))
|
||||
ranked = np.argsort(-scores.reshape(n, -1), axis=1)
|
||||
result = np.zeros_like(base) if keep else base.copy()
|
||||
for row in range(n):
|
||||
for flat_index in ranked[row, :count]:
|
||||
modality, block = divmod(int(flat_index), N_BLOCKS)
|
||||
left, right = block * BLOCK, (block + 1) * BLOCK
|
||||
if keep:
|
||||
result[row, left:right, modality] = base[row, left:right, modality]
|
||||
else:
|
||||
result[row, left:right, modality] = False
|
||||
return result
|
||||
|
||||
|
||||
def _faithfulness_curves(
|
||||
model: AlignedFusionModel,
|
||||
split: Split,
|
||||
base_masks: np.ndarray,
|
||||
explanations: dict[str, tuple[np.ndarray, np.ndarray]],
|
||||
device: torch.device,
|
||||
seed: int,
|
||||
) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
|
||||
logits, pred_class, full_prob, full_reg = _selected_scores(model, split, base_masks, device)
|
||||
rng = np.random.default_rng(seed)
|
||||
random_cls = rng.random((split.n, 3, N_BLOCKS), dtype=np.float32)
|
||||
random_reg = random_cls.copy()
|
||||
curve_rows: list[dict[str, Any]] = []
|
||||
for method, (class_scores, reg_scores) in [*explanations.items(), ("random", (random_cls, random_reg))]:
|
||||
for target, scores in (("predicted_class_probability", class_scores), ("intensity", reg_scores)):
|
||||
for fraction in (0.10, 0.20, 0.30, 0.40, 0.50):
|
||||
delete_masks = _rank_delete_masks(base_masks, scores, fraction, keep=False)
|
||||
_, _, after_prob, after_reg = _selected_scores(model, split, delete_masks, device,
|
||||
target_classes=pred_class)
|
||||
if target == "predicted_class_probability":
|
||||
difference = full_prob - after_prob
|
||||
abs_difference = np.abs(difference)
|
||||
else:
|
||||
difference = full_reg - after_reg
|
||||
abs_difference = np.abs(difference)
|
||||
curve_rows.append({
|
||||
"method": method, "target": target, "fraction_removed": fraction,
|
||||
"mean_signed_drop": float(np.mean(difference)),
|
||||
"mean_absolute_change": float(np.mean(abs_difference)),
|
||||
"n_valid": split.n,
|
||||
})
|
||||
keep_masks = _rank_delete_masks(base_masks, scores, 0.30, keep=True)
|
||||
_, _, keep_prob, keep_reg = _selected_scores(model, split, keep_masks, device,
|
||||
target_classes=pred_class)
|
||||
if target == "predicted_class_probability":
|
||||
sufficiency = np.abs(full_prob - keep_prob)
|
||||
else:
|
||||
sufficiency = np.abs(full_reg - keep_reg)
|
||||
curve_rows.append({
|
||||
"method": method, "target": target, "fraction_removed": -0.30,
|
||||
"mean_signed_drop": float(np.mean(sufficiency)),
|
||||
"mean_absolute_change": float(np.mean(sufficiency)),
|
||||
"n_valid": split.n,
|
||||
})
|
||||
|
||||
summary_rows: list[dict[str, Any]] = []
|
||||
for method in explanations.keys() | {"random"}:
|
||||
for target in ("predicted_class_probability", "intensity"):
|
||||
local = [r for r in curve_rows if r["method"] == method and r["target"] == target]
|
||||
removal30 = next(r for r in local if r["fraction_removed"] == 0.30)
|
||||
sufficiency = next(r for r in local if r["fraction_removed"] == -0.30)
|
||||
removal = [r for r in local if r["fraction_removed"] > 0]
|
||||
auc = float(np.trapezoid([r["mean_signed_drop"] for r in removal], [r["fraction_removed"] for r in removal]))
|
||||
summary_rows.append({
|
||||
"method": method,
|
||||
"target": target,
|
||||
"comprehensiveness_signed_drop_at_30": removal30["mean_signed_drop"],
|
||||
"absolute_prediction_change_at_30": removal30["mean_absolute_change"],
|
||||
"sufficiency_abs_error_top_30": sufficiency["mean_absolute_change"],
|
||||
"deletion_drop_auc_10_to_50": auc,
|
||||
})
|
||||
return curve_rows, summary_rows
|
||||
|
||||
|
||||
def _noise_split(split: Split, seed: int, sigma: float = 0.02) -> Split:
|
||||
rng = np.random.default_rng(seed)
|
||||
xs = []
|
||||
for modality, x in enumerate(split.x):
|
||||
noise = rng.normal(0.0, sigma, size=x.shape).astype(np.float32)
|
||||
noise *= split.mask[:, :, modality, None]
|
||||
xs.append((x + noise).astype(np.float32))
|
||||
return Split(tuple(xs), split.mask.copy(), split.y_cls, split.y_reg, split.ids)
|
||||
|
||||
|
||||
def _balanced_subset(split: Split, count: int, seed: int) -> Split:
|
||||
rng = np.random.default_rng(seed)
|
||||
selected: list[int] = []
|
||||
per_class = max(1, count // 3)
|
||||
for label in (0, 1, 2):
|
||||
available = np.flatnonzero(split.y_cls == label)
|
||||
take = min(per_class, len(available))
|
||||
selected.extend(rng.choice(available, size=take, replace=False).tolist())
|
||||
if len(selected) < count:
|
||||
remaining = np.setdiff1d(np.arange(split.n), np.asarray(selected, dtype=int))
|
||||
extra = min(count - len(selected), len(remaining))
|
||||
selected.extend(rng.choice(remaining, size=extra, replace=False).tolist())
|
||||
ids = np.asarray(sorted(selected[:count]), dtype=int)
|
||||
return Split(tuple(x[ids] for x in split.x), split.mask[ids], split.y_cls[ids], split.y_reg[ids], [split.ids[i] for i in ids])
|
||||
|
||||
|
||||
def _rank_stability(original: np.ndarray, changed: np.ndarray) -> tuple[float, float]:
|
||||
correlations, overlaps = [], []
|
||||
n, modalities, blocks = original.shape
|
||||
top_n = max(1, int(round(modalities * blocks * 0.30)))
|
||||
for row in range(n):
|
||||
a = original[row].reshape(-1)
|
||||
b = changed[row].reshape(-1)
|
||||
corr = spearmanr(a, b).statistic
|
||||
correlations.append(float(corr) if np.isfinite(corr) else 0.0)
|
||||
top_a = set(np.argsort(-a)[:top_n].tolist())
|
||||
top_b = set(np.argsort(-b)[:top_n].tolist())
|
||||
overlaps.append(len(top_a & top_b) / max(1, len(top_a | top_b)))
|
||||
return float(np.mean(correlations)), float(np.mean(overlaps))
|
||||
|
||||
|
||||
def _plot_faithfulness(curves: list[dict[str, Any]], output: Path) -> None:
|
||||
fig, axes = plt.subplots(1, 2, figsize=(10, 4.1), constrained_layout=True)
|
||||
styles = {"integrated_gradients": "#4e79a7", "grouped_occlusion": "#f28e2b", "random": "#999999"}
|
||||
for ax, target, title, ylabel in (
|
||||
(axes[0], "predicted_class_probability", "Polarity evidence deletion", "probability drop"),
|
||||
(axes[1], "intensity", "Intensity evidence deletion", "absolute intensity change"),
|
||||
):
|
||||
for method in styles:
|
||||
rows = sorted([r for r in curves if r["target"] == target and r["method"] == method and r["fraction_removed"] > 0], key=lambda r: r["fraction_removed"])
|
||||
if rows:
|
||||
metric = "mean_signed_drop" if target == "predicted_class_probability" else "mean_absolute_change"
|
||||
ax.plot([r["fraction_removed"] for r in rows], [r[metric] for r in rows], marker="o", label=method, color=styles[method])
|
||||
ax.set(title=title, xlabel="top evidence blocks removed", ylabel=ylabel)
|
||||
ax.grid(alpha=0.25)
|
||||
ax.legend(frameon=False)
|
||||
output.parent.mkdir(parents=True, exist_ok=True)
|
||||
fig.savefig(output, dpi=180)
|
||||
plt.close(fig)
|
||||
|
||||
|
||||
def _word_spans(text: str) -> list[tuple[int, int, str]]:
|
||||
return [(m.start(), m.end(), m.group(0)) for m in re.finditer(r"\S+", text)]
|
||||
|
||||
|
||||
def _offset_to_word(offset: tuple[int, int], spans: list[tuple[int, int, str]]) -> int | None:
|
||||
start, end = int(offset[0]), int(offset[1])
|
||||
if end <= start:
|
||||
return None
|
||||
overlaps = [max(0, min(end, right) - max(start, left)) for left, right, _ in spans]
|
||||
if not overlaps or max(overlaps) == 0:
|
||||
return None
|
||||
return int(np.argmax(overlaps))
|
||||
|
||||
|
||||
def _attachment4_raw() -> tuple[list[dict[str, Any]], list[Path]]:
|
||||
records, videos = [], []
|
||||
pkl_paths = sorted(ATTACHMENT4.glob("*.pkl"))
|
||||
for path in pkl_paths:
|
||||
with path.open("rb") as stream:
|
||||
record = pickle.load(stream, encoding="latin1")
|
||||
records.append(record)
|
||||
videos.append(ATTACHMENT4 / "videos" / f"{record['id']}.mp4")
|
||||
if len(records) != 20:
|
||||
raise ValueError(f"expected 20 aligned Attachment 4 clips; found {len(records)} in {ATTACHMENT4}")
|
||||
return records, videos
|
||||
|
||||
|
||||
def _attachment4_split(records: list[dict[str, Any]]) -> Split:
|
||||
xs = [[], [], []]
|
||||
masks = []
|
||||
ids = []
|
||||
for record in records:
|
||||
xs[0].append(np.asarray(record["text"], dtype=np.float32))
|
||||
xs[1].append(np.asarray(record["audio"], dtype=np.float32))
|
||||
xs[2].append(np.asarray(record["vision"], dtype=np.float32))
|
||||
token = np.asarray(record["text_bert"])
|
||||
masks.append(np.stack((token[1].astype(bool), np.any(record["audio"] != 0, axis=-1), np.any(record["vision"] != 0, axis=-1)), axis=-1))
|
||||
ids.append(str(record["id"]))
|
||||
return Split(tuple(np.stack(x) for x in xs), np.stack(masks), np.zeros(len(records), dtype=np.int64), np.zeros(len(records), dtype=np.float32), ids)
|
||||
|
||||
|
||||
def _block_value_per_slot(group_scores: np.ndarray, masks: np.ndarray) -> np.ndarray:
|
||||
n = group_scores.shape[0]
|
||||
slots = np.zeros((n, 3, 50), dtype=np.float32)
|
||||
for modality in range(3):
|
||||
for block in range(N_BLOCKS):
|
||||
left, right = block * BLOCK, (block + 1) * BLOCK
|
||||
active = masks[:, left:right, modality]
|
||||
count = active.sum(axis=1).clip(min=1)
|
||||
each = group_scores[:, modality, block] / count
|
||||
slots[:, modality, left:right] = each[:, None]
|
||||
return slots
|
||||
|
||||
|
||||
def _run_attachment4(
|
||||
model: AlignedFusionModel,
|
||||
output: Path,
|
||||
stats: RobustStats,
|
||||
device: torch.device,
|
||||
bert_tokenizer,
|
||||
class_method: str,
|
||||
reg_method: str,
|
||||
) -> None:
|
||||
records, videos = _attachment4_raw()
|
||||
raw = _attachment4_split(records)
|
||||
split = apply_robust_stats(raw, stats)
|
||||
logits, pred_class, prob, intensity = _selected_scores(model, split, split.mask, device)
|
||||
need_ig = class_method == "integrated_gradients" or reg_method == "integrated_gradients"
|
||||
need_occ = class_method == "grouped_occlusion" or reg_method == "grouped_occlusion"
|
||||
ig_class, ig_reg = _integrated_groups(model, split, split.mask, device) if need_ig else (None, None)
|
||||
occ_class, occ_reg = _occlusion_groups(model, split, split.mask, device) if need_occ else (None, None)
|
||||
class_group = ig_class if class_method == "integrated_gradients" else occ_class
|
||||
reg_group = ig_reg if reg_method == "integrated_gradients" else occ_reg
|
||||
|
||||
predictions: list[dict[str, Any]] = []
|
||||
evidence: list[dict[str, Any]] = []
|
||||
word_mappings: dict[str, list[dict[str, Any]]] = {}
|
||||
ctc_word_coverages: list[float] = []
|
||||
ctc_tokenizer, ctc_model = load_ctc(device)
|
||||
for index, (record, video_path) in enumerate(zip(records, videos)):
|
||||
clip_id = str(record["id"])
|
||||
text = str(record["raw_text"])
|
||||
words = text.split()
|
||||
time_status = "ok"
|
||||
try:
|
||||
waveform = decode_audio(video_path)
|
||||
intervals = align_words(waveform, words, ctc_tokenizer, ctc_model, device)
|
||||
except Exception as exc:
|
||||
intervals = []
|
||||
time_status = f"ctc_failed:{type(exc).__name__}"
|
||||
valid_word_count = sum(interval.valid for interval in intervals)
|
||||
ctc_coverage = valid_word_count / max(1, len(words))
|
||||
ctc_word_coverages.append(ctc_coverage)
|
||||
if time_status == "ok":
|
||||
time_status = "ok" if ctc_coverage >= 0.95 else ("partial" if valid_word_count else "failed")
|
||||
encoded = bert_tokenizer(text, padding="max_length", truncation=True, max_length=50,
|
||||
return_offsets_mapping=True, return_tensors="np")
|
||||
offsets = encoded["offset_mapping"][0]
|
||||
model_tokens = np.asarray(record["text_bert"])[0]
|
||||
input_ids_match = bool(np.array_equal(encoded["input_ids"][0], model_tokens))
|
||||
pieces = bert_tokenizer.convert_ids_to_tokens(model_tokens.tolist())
|
||||
spans = _word_spans(text)
|
||||
token_word = [_offset_to_word(tuple(offsets[i]), spans) for i in range(50)]
|
||||
local_words = []
|
||||
for slot in range(50):
|
||||
word_index = token_word[slot]
|
||||
interval = intervals[word_index] if word_index is not None and word_index < len(intervals) else None
|
||||
local_words.append({
|
||||
"slot": slot,
|
||||
"token": pieces[slot],
|
||||
"word_index": word_index,
|
||||
"word": spans[word_index][2] if word_index is not None else "",
|
||||
"start_s": interval.start_s if interval and interval.valid else float("nan"),
|
||||
"end_s": interval.end_s if interval and interval.valid else float("nan"),
|
||||
"ctc_quality": interval.quality if interval and interval.valid else 0.0,
|
||||
"ctc_valid": bool(interval and interval.valid),
|
||||
})
|
||||
word_mappings[clip_id] = local_words
|
||||
|
||||
predictions.append({
|
||||
"sample_id": clip_id,
|
||||
"predicted_class": CLASS_NAMES[int(pred_class[index])],
|
||||
"predicted_class_id": int(pred_class[index]),
|
||||
"predicted_class_probability": float(prob[index]),
|
||||
"predicted_intensity": float(intensity[index]),
|
||||
"transcript": text,
|
||||
"video_file_exists": video_path.is_file(),
|
||||
"ctc_alignment_status": time_status,
|
||||
"ctc_word_coverage": ctc_coverage,
|
||||
"ctc_aligned_words": valid_word_count,
|
||||
"transcript_words": len(words),
|
||||
"bert_token_ids_match_pickle": input_ids_match,
|
||||
})
|
||||
|
||||
for modality in range(3):
|
||||
block_values = class_group[index, modality]
|
||||
available_blocks = [
|
||||
block for block in range(N_BLOCKS)
|
||||
if np.any(split.mask[index, block * BLOCK:(block + 1) * BLOCK, modality])
|
||||
]
|
||||
top_blocks = sorted(available_blocks, key=lambda block: -float(block_values[block]))[:3]
|
||||
for block in top_blocks:
|
||||
left, right = block * BLOCK, (block + 1) * BLOCK
|
||||
local_slots = [slot for slot in range(left, right)
|
||||
if split.mask[index, slot, modality] and local_words[slot]["ctc_valid"]]
|
||||
if not local_slots:
|
||||
continue
|
||||
maps = [local_words[slot] for slot in local_slots]
|
||||
word_rows = {}
|
||||
for mapping in maps:
|
||||
if mapping["word_index"] is not None:
|
||||
word_rows[int(mapping["word_index"])] = mapping
|
||||
unique_words = [word_rows[key] for key in sorted(word_rows)]
|
||||
if not unique_words:
|
||||
continue
|
||||
evidence.append({
|
||||
"sample_id": clip_id,
|
||||
"modality": MODALITY_LABELS[modality],
|
||||
"block_index": int(block),
|
||||
"slot_start_index": int(left),
|
||||
"slot_end_index_exclusive": int(right),
|
||||
"slot_indices": ",".join(str(slot) for slot in local_slots),
|
||||
"tokens_or_wordpieces": " ".join(local_words[slot]["token"] for slot in local_slots),
|
||||
"matched_words": " ".join(mapping["word"] for mapping in unique_words),
|
||||
"word_indices": ",".join(str(mapping["word_index"]) for mapping in unique_words),
|
||||
"time_start_s": min(mapping["start_s"] for mapping in unique_words),
|
||||
"time_end_s": max(mapping["end_s"] for mapping in unique_words),
|
||||
"ctc_quality_uncalibrated_mean": float(np.mean([mapping["ctc_quality"] for mapping in unique_words])),
|
||||
"ctc_words_covered": len(unique_words),
|
||||
"ctc_word_coverage_clip": ctc_coverage,
|
||||
"class_importance": float(class_group[index, modality, block]),
|
||||
"intensity_importance": float(reg_group[index, modality, block]),
|
||||
"class_explainer": class_method,
|
||||
"intensity_explainer": reg_method,
|
||||
"ctc_alignment_status": time_status,
|
||||
})
|
||||
|
||||
_write_csv(output / "attachment4_predictions.csv", predictions)
|
||||
_write_csv(output / "attachment4_top_evidence.csv", evidence)
|
||||
_plot_attachment4_example(records, videos, predictions, evidence, output)
|
||||
(output / "attachment4_alignment_audit.json").write_text(json.dumps({
|
||||
"n_samples": len(records),
|
||||
"n_video_files_found": sum(x.is_file() for x in videos),
|
||||
"n_ctc_any_words_aligned": sum(row["ctc_aligned_words"] > 0 for row in predictions),
|
||||
"n_ctc_full_word_coverage": sum(row["ctc_alignment_status"] == "ok" for row in predictions),
|
||||
"mean_transcript_word_coverage": float(np.mean(ctc_word_coverages)),
|
||||
"n_bert_token_sequences_matching_pickle": sum(row["bert_token_ids_match_pickle"] for row in predictions),
|
||||
"time_mapping": "Q1 B1 CTC Viterbi hard word intervals computed from the supplied Attachment 4 video audio and transcript; subword slots inherit their transcript word interval",
|
||||
"quality_note": "CTC path score is uncalibrated. These intervals are localization references for interpretation, not human-annotated ground truth.",
|
||||
}, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
|
||||
|
||||
def _plot_attachment4_example(records, videos, predictions, evidence, output: Path) -> None:
|
||||
eligible = [row for row in predictions if row["ctc_alignment_status"] in {"ok", "partial"}]
|
||||
if not eligible:
|
||||
return
|
||||
chosen = eligible[0]
|
||||
clip_id = chosen["sample_id"]
|
||||
transcript = str(next(r["raw_text"] for r in records if str(r["id"]) == clip_id))
|
||||
local = [row for row in evidence if row["sample_id"] == clip_id
|
||||
and float(row["time_end_s"]) > float(row["time_start_s"])]
|
||||
if not local:
|
||||
return
|
||||
word_salience: dict[tuple[str, int, str], float] = {}
|
||||
word_times: dict[tuple[str, int, str], tuple[float, float, str]] = {}
|
||||
for row in local:
|
||||
key = (row["modality"], int(row["block_index"]), row["matched_words"])
|
||||
word_salience[key] = word_salience.get(key, 0.0) + float(row["class_importance"])
|
||||
word_times[key] = (float(row["time_start_s"]), float(row["time_end_s"]), row["matched_words"])
|
||||
if not word_times:
|
||||
return
|
||||
max_time = max(value[1] for value in word_times.values())
|
||||
fig, ax = plt.subplots(figsize=(12, 4.2), constrained_layout=True)
|
||||
palette = {"text": "#4e79a7", "audio": "#f28e2b", "vision": "#59a14f"}
|
||||
max_value = max(word_salience.values(), default=1.0) or 1.0
|
||||
y_levels = {"text": 2, "audio": 1, "vision": 0}
|
||||
for key, salience in word_salience.items():
|
||||
modality, _, word = key
|
||||
if key not in word_times:
|
||||
continue
|
||||
start, end, _ = word_times[key]
|
||||
alpha = 0.25 + 0.75 * min(1.0, salience / max_value)
|
||||
y = y_levels[modality]
|
||||
ax.broken_barh([(start, max(0.01, end - start))], (y - 0.3, 0.6),
|
||||
facecolors=palette[modality], alpha=alpha, edgecolors="white", linewidth=0.35)
|
||||
ax.text((start + end) / 2, y, word, ha="center", va="center", fontsize=6, rotation=55)
|
||||
ax.set_yticks([0, 1, 2], labels=["Vision", "Audio", "Text"])
|
||||
ax.set_xlim(0, max(0.1, max_time))
|
||||
ax.set_xlabel("seconds from clip start (Q1 CTC word-time mapping)")
|
||||
ax.set_title(f"Attachment 4 example {clip_id}: {chosen['predicted_class']} / intensity {chosen['predicted_intensity']:.2f}")
|
||||
ax.grid(axis="x", alpha=0.2)
|
||||
output.mkdir(parents=True, exist_ok=True)
|
||||
fig.savefig(output / f"attachment4_{clip_id}_evidence_timeline.png", dpi=180)
|
||||
plt.close(fig)
|
||||
|
||||
|
||||
def _run(args: argparse.Namespace) -> None:
|
||||
output = Path(args.output_dir)
|
||||
output.mkdir(parents=True, exist_ok=True)
|
||||
q2_output = Path(args.q2_output_dir)
|
||||
device = torch.device("cuda" if args.device == "auto" and torch.cuda.is_available() else ("cpu" if args.device == "auto" else args.device))
|
||||
torch.set_num_threads(args.threads)
|
||||
method, model = _model_from_run(q2_output, device)
|
||||
stats = RobustStats.load(q2_output / "aligned_robust_stats.npz")
|
||||
valid = apply_robust_stats(load_aligned()["valid"], stats)
|
||||
|
||||
started = time.perf_counter()
|
||||
ig_class, ig_reg = _integrated_groups(model, valid, valid.mask, device, steps=args.ig_steps, batch_size=args.batch_size)
|
||||
ig_seconds = time.perf_counter() - started
|
||||
started = time.perf_counter()
|
||||
occ_class, occ_reg = _occlusion_groups(model, valid, valid.mask, device)
|
||||
occ_seconds = time.perf_counter() - started
|
||||
explanations = {"integrated_gradients": (ig_class, ig_reg), "grouped_occlusion": (occ_class, occ_reg)}
|
||||
curves, summary = _faithfulness_curves(model, valid, valid.mask, explanations, device, args.seed)
|
||||
|
||||
stability_subset = _balanced_subset(valid, args.stability_samples, args.seed + 17)
|
||||
noisy_subset = _noise_split(stability_subset, args.seed + 23, sigma=args.noise_sigma)
|
||||
stable_ig_class, stable_ig_reg = _integrated_groups(model, noisy_subset, noisy_subset.mask, device,
|
||||
steps=args.ig_steps, batch_size=args.batch_size)
|
||||
stable_occ_class, stable_occ_reg = _occlusion_groups(model, noisy_subset, noisy_subset.mask, device)
|
||||
stability_rows = []
|
||||
for name, original, perturbed in (
|
||||
("integrated_gradients", ig_class[np.isin(np.asarray(valid.ids), stability_subset.ids)], stable_ig_class),
|
||||
("grouped_occlusion", occ_class[np.isin(np.asarray(valid.ids), stability_subset.ids)], stable_occ_class),
|
||||
):
|
||||
corr, jaccard = _rank_stability(original, perturbed)
|
||||
stability_rows.append({"method": name, "target": "predicted_class_probability", "spearman_rank_correlation": corr,
|
||||
"top_30_percent_jaccard": jaccard, "n_samples": len(stability_subset.ids),
|
||||
"input_noise_sigma": args.noise_sigma})
|
||||
for name, original, perturbed in (
|
||||
("integrated_gradients", ig_reg[np.isin(np.asarray(valid.ids), stability_subset.ids)], stable_ig_reg),
|
||||
("grouped_occlusion", occ_reg[np.isin(np.asarray(valid.ids), stability_subset.ids)], stable_occ_reg),
|
||||
):
|
||||
corr, jaccard = _rank_stability(original, perturbed)
|
||||
stability_rows.append({"method": name, "target": "intensity", "spearman_rank_correlation": corr,
|
||||
"top_30_percent_jaccard": jaccard, "n_samples": len(stability_subset.ids),
|
||||
"input_noise_sigma": args.noise_sigma})
|
||||
|
||||
for row in summary:
|
||||
row["runtime_seconds"] = ig_seconds if row["method"] == "integrated_gradients" else occ_seconds
|
||||
stability = next((s for s in stability_rows if s["method"] == row["method"] and s["target"] == row["target"]), None)
|
||||
if stability:
|
||||
row.update(stability)
|
||||
else:
|
||||
row.update({"spearman_rank_correlation": float("nan"), "top_30_percent_jaccard": float("nan")})
|
||||
_write_csv(output / "q3_explanation_method_summary.csv", summary)
|
||||
_write_csv(output / "q3_deletion_curves.csv", curves)
|
||||
_write_csv(output / "q3_explanation_stability.csv", stability_rows)
|
||||
_plot_faithfulness(curves, output / "q3_explanation_faithfulness.png")
|
||||
|
||||
# Select classification and intensity explainers separately, based on direct validation probes.
|
||||
cls = [r for r in summary if r["target"] == "predicted_class_probability" and r["method"] != "random"]
|
||||
reg = [r for r in summary if r["target"] == "intensity" and r["method"] != "random"]
|
||||
class_method = sorted(cls, key=lambda r: (-r["comprehensiveness_signed_drop_at_30"], r["sufficiency_abs_error_top_30"], r["method"]))[0]["method"]
|
||||
reg_ig = next(r for r in reg if r["method"] == "integrated_gradients")
|
||||
reg_occ = next(r for r in reg if r["method"] == "grouped_occlusion")
|
||||
# There is a real tradeoff for the regression head: IG changes the output more
|
||||
# after deletion, while occlusion better retains it when only the selected
|
||||
# evidence is kept. Use direct intervention for the displayed segments and
|
||||
# retain IG as a directional cross-check.
|
||||
reg_method = "grouped_occlusion"
|
||||
selection = {
|
||||
"q2_predictor": method,
|
||||
"classification_explainer": class_method,
|
||||
"intensity_explainer_primary": reg_method,
|
||||
"intensity_explainer_crosscheck": "integrated_gradients",
|
||||
"intensity_tradeoff": {
|
||||
"integrated_gradients_abs_change_at_30": reg_ig["absolute_prediction_change_at_30"],
|
||||
"integrated_gradients_sufficiency_error_top_30": reg_ig["sufficiency_abs_error_top_30"],
|
||||
"grouped_occlusion_abs_change_at_30": reg_occ["absolute_prediction_change_at_30"],
|
||||
"grouped_occlusion_sufficiency_error_top_30": reg_occ["sufficiency_abs_error_top_30"],
|
||||
},
|
||||
"selection_basis": "For polarity, grouped occlusion has the larger signed target-probability drop, lower sufficiency error, higher deletion AUC, and lower runtime. For intensity, IG causes a larger deletion change but grouped occlusion has lower top-evidence sufficiency error; grouped occlusion is used for displayed segments and IG is retained as a cross-check. No combined explanation score is used.",
|
||||
"valid_samples": valid.n,
|
||||
"stability_samples": len(stability_subset.ids),
|
||||
"integrated_gradients_runtime_seconds": ig_seconds,
|
||||
"grouped_occlusion_runtime_seconds": occ_seconds,
|
||||
"ctc_time_map_for_attachment4": "Q1 hard CTC Viterbi word boundaries from source video audio; not human alignment ground truth",
|
||||
}
|
||||
(output / "q3_explainer_selection.json").write_text(json.dumps(selection, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
print(f"Q3 explainers: classification={class_method}; intensity={reg_method}", flush=True)
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained("google-bert/bert-base-uncased", use_fast=True)
|
||||
_run_attachment4(model, output, stats, device, tokenizer, class_method, reg_method)
|
||||
print(f"saved Q3 explanation selection and Attachment 4 evidence to {output}", flush=True)
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser(description="Compare faithful Q3 explanations and map Attachment 4 evidence to video time")
|
||||
parser.add_argument("--q2-output-dir", default=str(Q2_PROJECT / "outputs" / "algorithm_selection"))
|
||||
parser.add_argument("--output-dir", default=str(Path(__file__).resolve().parents[1] / "outputs" / "explanation_selection"))
|
||||
parser.add_argument("--device", default="auto")
|
||||
parser.add_argument("--threads", type=int, default=4)
|
||||
parser.add_argument("--batch-size", type=int, default=48)
|
||||
parser.add_argument("--ig-steps", type=int, default=16)
|
||||
parser.add_argument("--stability-samples", type=int, default=120)
|
||||
parser.add_argument("--noise-sigma", type=float, default=0.02)
|
||||
parser.add_argument("--seed", type=int, default=42)
|
||||
args = parser.parse_args()
|
||||
_run(args)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user