Add Q3 MoFE router visualizations and explanations
This commit is contained in:
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 01
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **neutral**
|
||||
- Intensity: -0.172
|
||||
- Predicted-class confidence: 0.565
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/01.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.0003 | 0.819 | -0.4419 | 0.510 |
|
||||
| audio | -0.1551 | 0.127 | +0.1493 | 0.172 |
|
||||
| vision | -0.0656 | 0.054 | -0.2749 | 0.317 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0683 | -0.0865 |
|
||||
| Text + Vision | -0.0051 | +0.3554 |
|
||||
| Audio + Vision | +0.0441 | +0.1078 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 19–21 (relative progress 0.380–0.440), estimated clip interval 3.42–3.96s** — supports_predicted_class; evidence: belt is
|
||||
- **text bins 47–47 (relative progress 0.940–0.960), estimated clip interval 8.46–8.64s** — supports_predicted_class; evidence: requirements
|
||||
- **audio bins 22–24 (relative progress 0.440–0.500), estimated clip interval 3.96–4.50s** — opposes_predicted_class; evidence: unaligned audio feature rows 78–89; review the linked source clip at the estimated relative span
|
||||
- **audio bins 3–4 (relative progress 0.060–0.100), estimated clip interval 0.54–0.90s** — opposes_predicted_class; evidence: unaligned audio feature rows 10–17; review the linked source clip at the estimated relative span
|
||||
- **vision bins 46–48 (relative progress 0.920–0.980), estimated clip interval 8.28–8.82s** — opposes_predicted_class; evidence: unaligned vision feature rows 124–132; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 32–33 (relative progress 0.640–0.680), estimated clip interval 5.76–6.12s** — opposes_predicted_class; evidence: unaligned vision feature rows 86–91; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3528 / +0.9214 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.1364 / -0.3016. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.7019.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Replacing these wear components when replacing the timing belt is essential to ensuring the new belt performs to its mileage requirements
|
||||
@@ -0,0 +1,56 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 02
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.395
|
||||
- Predicted-class confidence: 0.634
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/02.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +0.6519 | 0.573 | +0.1487 | 0.332 |
|
||||
| audio | +0.1355 | 0.119 | +0.0748 | 0.167 |
|
||||
| vision | -0.3508 | 0.308 | -0.2239 | 0.500 |
|
||||
|
||||
Shapley completeness residuals: class 5.55e-17, intensity -2.78e-17.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0716 | +0.0491 |
|
||||
| Text + Vision | +0.1147 | +0.0187 |
|
||||
| Audio + Vision | +0.0890 | +0.0524 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 22–26 (relative progress 0.440–0.540), estimated clip interval 1.54–1.89s** — supports_predicted_class; evidence: s happiness -
|
||||
- **audio bins 35–39 (relative progress 0.700–0.800), estimated clip interval 2.45–2.79s** — supports_predicted_class; evidence: unaligned audio feature rows 46–52; review the linked source clip at the estimated relative span
|
||||
- **vision bins 10–12 (relative progress 0.200–0.260), estimated clip interval 0.70–0.91s** — opposes_predicted_class; evidence: unaligned vision feature rows 10–12; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 34–34 (relative progress 0.680–0.700), estimated clip interval 2.38–2.45s** — opposes_predicted_class; evidence: unaligned vision feature rows 34–35; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.4557 / +0.4939 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -1.1868 / -0.5516. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.4518.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
We want to live by each other’s happiness - not by each other’s misery.
|
||||
@@ -0,0 +1,57 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 03
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.493
|
||||
- Predicted-class confidence: 0.553
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/03.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.1502 | 0.628 | -1.0738 | 0.671 |
|
||||
| audio | +0.1589 | 0.087 | -0.1704 | 0.106 |
|
||||
| vision | -0.5213 | 0.285 | +0.3565 | 0.223 |
|
||||
|
||||
Shapley completeness residuals: class -1.11e-16, intensity 2.22e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.1032 | +0.2050 |
|
||||
| Text + Vision | +0.3268 | -0.2207 |
|
||||
| Audio + Vision | +0.0454 | +0.0051 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 30–34 (relative progress 0.600–0.700), estimated clip interval 5.80–6.76s** — supports_predicted_class; evidence: take rental income into
|
||||
- **audio bins 14–17 (relative progress 0.280–0.360), estimated clip interval 2.71–3.48s** — supports_predicted_class; evidence: unaligned audio feature rows 53–68; review the linked source clip at the estimated relative span
|
||||
- **audio bins 48–48 (relative progress 0.960–0.980), estimated clip interval 9.28–9.47s** — supports_predicted_class; evidence: unaligned audio feature rows 183–187; review the linked source clip at the estimated relative span
|
||||
- **vision bins 17–19 (relative progress 0.340–0.400), estimated clip interval 3.29–3.87s** — opposes_predicted_class; evidence: unaligned vision feature rows 48–57; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 2–2 (relative progress 0.040–0.060), estimated clip interval 0.39–0.58s** — opposes_predicted_class; evidence: unaligned vision feature rows 5–8; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.6246 / +1.3268 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -1.1954 / -0.3649. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.0492.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
There's one lender at the moment which I think is just Bankwest who don't take rental income into account, they take rental yield into account.
|
||||
@@ -0,0 +1,56 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 04
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.968
|
||||
- Predicted-class confidence: 0.769
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/04.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.6860 | 0.849 | -1.7115 | 0.831 |
|
||||
| audio | -0.2440 | 0.123 | +0.2778 | 0.135 |
|
||||
| vision | +0.0555 | 0.028 | +0.0707 | 0.034 |
|
||||
|
||||
Shapley completeness residuals: class -4.44e-16, intensity 4.44e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.1306 | -0.0822 |
|
||||
| Text + Vision | -0.0863 | +0.0593 |
|
||||
| Audio + Vision | -0.0342 | +0.0247 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 5–6 (relative progress 0.100–0.140), estimated clip interval 0.70–0.98s** — supports_predicted_class; evidence: i blow
|
||||
- **text bins 44–45 (relative progress 0.880–0.920), estimated clip interval 6.15–6.43s** — supports_predicted_class; evidence: absolutely not
|
||||
- **audio bins 11–14 (relative progress 0.220–0.300), estimated clip interval 1.54–2.10s** — opposes_predicted_class; evidence: unaligned audio feature rows 30–41; review the linked source clip at the estimated relative span
|
||||
- **audio bins 37–37 (relative progress 0.740–0.760), estimated clip interval 5.18–5.32s** — opposes_predicted_class; evidence: unaligned audio feature rows 101–104; review the linked source clip at the estimated relative span
|
||||
- **vision bins 35–39 (relative progress 0.700–0.800), estimated clip interval 4.90–5.60s** — opposes_predicted_class; evidence: unaligned vision feature rows 72–82; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.9337 / +1.7442 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.6584 / -0.3517. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.4210.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
If I blow it at the team exercise, should I kiss my chances of cheering "GO BLUE" goodbye?] Absolutely not.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 05
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +1.129
|
||||
- Predicted-class confidence: 0.848
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/05.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +0.4992 | 0.376 | +0.0947 | 0.129 |
|
||||
| audio | +0.1165 | 0.088 | +0.0709 | 0.097 |
|
||||
| vision | +0.7112 | 0.536 | +0.5687 | 0.774 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.1349 | +0.1991 |
|
||||
| Text + Vision | -0.1658 | +0.1123 |
|
||||
| Audio + Vision | -0.0532 | -0.0378 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 0–2 (relative progress 0.000–0.060), estimated clip interval 0.00–0.37s** — supports_predicted_class; evidence: hi
|
||||
- **text bins 6–7 (relative progress 0.120–0.160), estimated clip interval 0.73–0.98s** — supports_predicted_class; evidence: ,
|
||||
- **audio bins 18–20 (relative progress 0.360–0.420), estimated clip interval 2.20–2.56s** — supports_predicted_class; evidence: unaligned audio feature rows 43–50; review the linked source clip at the estimated relative span
|
||||
- **audio bins 40–41 (relative progress 0.800–0.840), estimated clip interval 4.88–5.12s** — supports_predicted_class; evidence: unaligned audio feature rows 96–100; review the linked source clip at the estimated relative span
|
||||
- **vision bins 21–24 (relative progress 0.420–0.500), estimated clip interval 2.56–3.05s** — supports_predicted_class; evidence: unaligned vision feature rows 37–44; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 44–44 (relative progress 0.880–0.900), estimated clip interval 5.37–5.49s** — supports_predicted_class; evidence: unaligned vision feature rows 79–80; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.2997 / +0.7538 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.0442 / +0.0098. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.7515.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Hi, my name is Chloe, video marketer for Red Wagon Marketing.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 06
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +1.201
|
||||
- Predicted-class confidence: 0.849
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/06.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.7299 | 0.798 | +1.0189 | 0.827 |
|
||||
| audio | -0.1445 | 0.067 | -0.2116 | 0.172 |
|
||||
| vision | -0.2938 | 0.135 | -0.0010 | 0.001 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.1320 | -0.0408 |
|
||||
| Text + Vision | +0.0014 | -0.1078 |
|
||||
| Audio + Vision | +0.1490 | +0.0687 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 8–10 (relative progress 0.160–0.220), estimated clip interval 1.33–1.83s** — supports_predicted_class; evidence: a fan of dancing
|
||||
- **text bins 21–22 (relative progress 0.420–0.460), estimated clip interval 3.49–3.82s** — supports_predicted_class; evidence: to watch people
|
||||
- **audio bins 5–8 (relative progress 0.100–0.180), estimated clip interval 0.83–1.49s** — opposes_predicted_class; evidence: unaligned audio feature rows 16–29; review the linked source clip at the estimated relative span
|
||||
- **audio bins 38–38 (relative progress 0.760–0.780), estimated clip interval 6.31–6.48s** — opposes_predicted_class; evidence: unaligned audio feature rows 123–127; review the linked source clip at the estimated relative span
|
||||
- **vision bins 33–36 (relative progress 0.660–0.740), estimated clip interval 5.48–6.15s** — opposes_predicted_class; evidence: unaligned vision feature rows 81–91; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 40–40 (relative progress 0.800–0.820), estimated clip interval 6.64–6.81s** — opposes_predicted_class; evidence: unaligned vision feature rows 98–100; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.5652 / +1.4485 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.4569 / -0.5398. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.1669.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
If you're a fan of dancing in that sense, just like to watch people dance, see impressive dance moves then you might want to check out this movie solely for that
|
||||
@@ -0,0 +1,57 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 07
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.852
|
||||
- Predicted-class confidence: 0.808
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/07.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.2071 | 0.949 | +0.7027 | 0.741 |
|
||||
| audio | +0.0179 | 0.014 | -0.0444 | 0.047 |
|
||||
| vision | -0.0466 | 0.037 | -0.2015 | 0.212 |
|
||||
|
||||
Shapley completeness residuals: class -2.22e-16, intensity -1.11e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.0059 | +0.1040 |
|
||||
| Text + Vision | +0.1672 | +0.1852 |
|
||||
| Audio + Vision | +0.0742 | +0.1056 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 23–27 (relative progress 0.460–0.560), estimated clip interval 9.82–11.95s** — supports_predicted_class; evidence: the topics of growing the
|
||||
- **audio bins 1–3 (relative progress 0.020–0.080), estimated clip interval 0.43–1.71s** — supports_predicted_class; evidence: unaligned audio feature rows 8–33; review the linked source clip at the estimated relative span
|
||||
- **audio bins 11–12 (relative progress 0.220–0.260), estimated clip interval 4.70–5.55s** — supports_predicted_class; evidence: unaligned audio feature rows 93–110; review the linked source clip at the estimated relative span
|
||||
- **vision bins 10–13 (relative progress 0.200–0.280), estimated clip interval 4.27–5.98s** — supports_predicted_class; evidence: unaligned vision feature rows 63–89; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 4–4 (relative progress 0.080–0.100), estimated clip interval 1.71–2.13s** — supports_predicted_class; evidence: unaligned vision feature rows 25–31; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.5147 / +1.1770 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.5447 / -0.0666. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.0335.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
As Linn’s associate editor Michael Baadke reports in our November 28 issue, attendees “enthusiastically discussed the topics of growing the hobby, the future of stamp shows, and dealers and philatelic partnerships, along with ways the leading organizations involved in the stamp hobby can work together to make it succeed and grow
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 08
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.836
|
||||
- Predicted-class confidence: 0.779
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/08.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.6981 | 0.750 | +0.7934 | 0.693 |
|
||||
| audio | -0.0794 | 0.035 | -0.0434 | 0.038 |
|
||||
| vision | -0.4875 | 0.215 | -0.3086 | 0.269 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0518 | -0.0102 |
|
||||
| Text + Vision | +0.2869 | +0.1244 |
|
||||
| Audio + Vision | +0.1742 | +0.1245 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 38–40 (relative progress 0.760–0.820), estimated clip interval 3.87–4.17s** — supports_predicted_class; evidence: design grand
|
||||
- **text bins 20–20 (relative progress 0.400–0.420), estimated clip interval 2.03–2.14s** — supports_predicted_class; evidence: tonight
|
||||
- **audio bins 40–41 (relative progress 0.800–0.840), estimated clip interval 4.07–4.27s** — opposes_predicted_class; evidence: unaligned audio feature rows 78–82; review the linked source clip at the estimated relative span
|
||||
- **audio bins 11–12 (relative progress 0.220–0.260), estimated clip interval 1.12–1.32s** — supports_predicted_class; evidence: unaligned audio feature rows 21–25; review the linked source clip at the estimated relative span
|
||||
- **vision bins 41–43 (relative progress 0.820–0.880), estimated clip interval 4.17–4.48s** — opposes_predicted_class; evidence: unaligned vision feature rows 60–65; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 3–4 (relative progress 0.060–0.100), estimated clip interval 0.31–0.51s** — opposes_predicted_class; evidence: unaligned vision feature rows 4–7; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.6036 / +1.7624 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.4005 / -0.3419. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.3297.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
That brings us to tonight, the Universal Design Grand Challenge
|
||||
@@ -0,0 +1,54 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 09
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.750
|
||||
- Predicted-class confidence: 0.920
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/09.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +2.4909 | 0.906 | -2.1191 | 0.988 |
|
||||
| audio | -0.0254 | 0.009 | -0.0083 | 0.004 |
|
||||
| vision | -0.2337 | 0.085 | -0.0172 | 0.008 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 4.44e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0573 | +0.0615 |
|
||||
| Text + Vision | -0.0030 | +0.1784 |
|
||||
| Audio + Vision | -0.0220 | +0.1153 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 28–32 (relative progress 0.560–0.660), estimated clip interval 2.82–3.33s** — supports_predicted_class; evidence: at all,
|
||||
- **audio bins 28–32 (relative progress 0.560–0.660), estimated clip interval 2.82–3.33s** — opposes_predicted_class; evidence: unaligned audio feature rows 55–65; review the linked source clip at the estimated relative span
|
||||
- **vision bins 6–10 (relative progress 0.120–0.220), estimated clip interval 0.61–1.11s** — opposes_predicted_class; evidence: unaligned vision feature rows 8–16; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.7794 / +2.2603 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.2221 / -0.3355. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.7260.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
(uhh) I did not like this movie at all, I would not recommend it
|
||||
@@ -0,0 +1,57 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 10
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.954
|
||||
- Predicted-class confidence: 0.783
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/10.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.7536 | 0.822 | -1.6119 | 0.855 |
|
||||
| audio | -0.2526 | 0.118 | +0.2683 | 0.142 |
|
||||
| vision | +0.1275 | 0.060 | -0.0050 | 0.003 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 2.22e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.0873 | -0.0349 |
|
||||
| Text + Vision | -0.1980 | +0.1134 |
|
||||
| Audio + Vision | +0.0271 | +0.0146 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 42–46 (relative progress 0.840–0.940), estimated clip interval 8.35–9.34s** — supports_predicted_class; evidence: this one was pretty terrible
|
||||
- **audio bins 3–5 (relative progress 0.060–0.120), estimated clip interval 0.60–1.19s** — opposes_predicted_class; evidence: unaligned audio feature rows 11–23; review the linked source clip at the estimated relative span
|
||||
- **audio bins 22–23 (relative progress 0.440–0.480), estimated clip interval 4.37–4.77s** — opposes_predicted_class; evidence: unaligned audio feature rows 85–93; review the linked source clip at the estimated relative span
|
||||
- **vision bins 4–5 (relative progress 0.080–0.120), estimated clip interval 0.80–1.19s** — supports_predicted_class; evidence: unaligned vision feature rows 11–17; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 23–24 (relative progress 0.460–0.500), estimated clip interval 4.57–4.97s** — supports_predicted_class; evidence: unaligned vision feature rows 67–73; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.8593 / +1.6509 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.8881 / -0.3659. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.3542.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
(umm) And you know I really do like to see fluffy chick flicks sometimes so I'm not against that but this one was pretty terrible
|
||||
@@ -0,0 +1,56 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 11
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.469
|
||||
- Predicted-class confidence: 0.614
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/11.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +0.6469 | 0.458 | -0.6779 | 0.525 |
|
||||
| audio | +0.5563 | 0.394 | -0.4001 | 0.310 |
|
||||
| vision | -0.2098 | 0.148 | +0.2138 | 0.165 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.2481 | +0.2153 |
|
||||
| Text + Vision | +0.0912 | -0.0661 |
|
||||
| Audio + Vision | -0.0434 | +0.0212 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 8–12 (relative progress 0.160–0.260), estimated clip interval 0.93–1.51s** — supports_predicted_class; evidence: would be ashamed
|
||||
- **audio bins 32–36 (relative progress 0.640–0.740), estimated clip interval 3.71–4.29s** — supports_predicted_class; evidence: unaligned audio feature rows 72–84; review the linked source clip at the estimated relative span
|
||||
- **vision bins 21–24 (relative progress 0.420–0.500), estimated clip interval 2.44–2.90s** — opposes_predicted_class; evidence: unaligned vision feature rows 36–42; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 33–33 (relative progress 0.660–0.680), estimated clip interval 3.83–3.95s** — opposes_predicted_class; evidence: unaligned vision feature rows 56–58; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.7060 / +0.7600 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -1.4195 / -0.1219. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.7509.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
I would be ashamed to have made this film if I was a director
|
||||
@@ -0,0 +1,56 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 12
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.238
|
||||
- Predicted-class confidence: 0.848
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/12.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.3737 | 0.766 | -1.3134 | 0.805 |
|
||||
| audio | +0.0008 | 0.000 | -0.0352 | 0.022 |
|
||||
| vision | +0.4181 | 0.233 | -0.2839 | 0.174 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0123 | +0.0978 |
|
||||
| Text + Vision | -0.2289 | +0.2750 |
|
||||
| Audio + Vision | -0.1519 | +0.0960 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 26–28 (relative progress 0.520–0.580), estimated clip interval 6.77–7.55s** — supports_predicted_class; evidence: by fault of their
|
||||
- **text bins 45–45 (relative progress 0.900–0.920), estimated clip interval 11.71–11.97s** — supports_predicted_class; evidence: is very
|
||||
- **audio bins 8–10 (relative progress 0.160–0.220), estimated clip interval 2.08–2.86s** — opposes_predicted_class; evidence: unaligned audio feature rows 41–56; review the linked source clip at the estimated relative span
|
||||
- **audio bins 24–24 (relative progress 0.480–0.500), estimated clip interval 6.25–6.51s** — opposes_predicted_class; evidence: unaligned audio feature rows 123–128; review the linked source clip at the estimated relative span
|
||||
- **vision bins 3–7 (relative progress 0.060–0.160), estimated clip interval 0.78–2.08s** — supports_predicted_class; evidence: unaligned vision feature rows 11–30; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.5112 / +1.1841 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.1864 / -0.0623. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.9530.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Or worse, an individual previously had good credit, but usually by no fault of their own, or perhaps by fault of their own, they have let their credit sag, and credit scores is very low.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 13
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **neutral**
|
||||
- Intensity: -0.056
|
||||
- Predicted-class confidence: 0.617
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/13.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.3506 | 0.767 | -0.3893 | 0.399 |
|
||||
| audio | +0.0614 | 0.035 | -0.3239 | 0.332 |
|
||||
| vision | -0.3485 | 0.198 | +0.2621 | 0.269 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0766 | +0.3903 |
|
||||
| Text + Vision | +0.1221 | +0.0147 |
|
||||
| Audio + Vision | +0.0191 | +0.1027 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 30–32 (relative progress 0.600–0.660), estimated clip interval 5.21–5.73s** — supports_predicted_class; evidence: a relationship between
|
||||
- **text bins 26–27 (relative progress 0.520–0.560), estimated clip interval 4.51–4.86s** — supports_predicted_class; evidence: i can
|
||||
- **audio bins 24–25 (relative progress 0.480–0.520), estimated clip interval 4.17–4.51s** — supports_predicted_class; evidence: unaligned audio feature rows 83–89; review the linked source clip at the estimated relative span
|
||||
- **audio bins 20–21 (relative progress 0.400–0.440), estimated clip interval 3.47–3.82s** — opposes_predicted_class; evidence: unaligned audio feature rows 69–76; review the linked source clip at the estimated relative span
|
||||
- **vision bins 45–47 (relative progress 0.900–0.960), estimated clip interval 7.81–8.33s** — opposes_predicted_class; evidence: unaligned vision feature rows 16–17; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 42–43 (relative progress 0.840–0.880), estimated clip interval 7.29–7.64s** — opposes_predicted_class; evidence: unaligned vision feature rows 15–15; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.4948 / +1.2805 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.1418 / -0.2321. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.9508.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
For example, I could take a set of data and from that data, I can find a relationship between any two of the given factors or more.
|
||||
@@ -0,0 +1,57 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 14
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.767
|
||||
- Predicted-class confidence: 0.698
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/14.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | -0.7961 | 0.343 | -0.6639 | 0.391 |
|
||||
| audio | +0.3769 | 0.163 | +0.1741 | 0.102 |
|
||||
| vision | +1.1452 | 0.494 | +0.8621 | 0.507 |
|
||||
|
||||
Shapley completeness residuals: class -2.22e-16, intensity -1.11e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.0732 | +0.1047 |
|
||||
| Text + Vision | -0.2140 | -0.0911 |
|
||||
| Audio + Vision | -0.1629 | -0.0936 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 33–35 (relative progress 0.660–0.720), estimated clip interval 6.83–7.45s** — opposes_predicted_class; evidence: ( ufe
|
||||
- **text bins 28–29 (relative progress 0.560–0.600), estimated clip interval 5.80–6.21s** — opposes_predicted_class; evidence: of uniform final
|
||||
- **audio bins 25–29 (relative progress 0.500–0.600), estimated clip interval 5.18–6.21s** — supports_predicted_class; evidence: unaligned audio feature rows 102–122; review the linked source clip at the estimated relative span
|
||||
- **vision bins 37–39 (relative progress 0.740–0.800), estimated clip interval 7.66–8.28s** — supports_predicted_class; evidence: unaligned vision feature rows 113–122; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 29–30 (relative progress 0.580–0.620), estimated clip interval 6.00–6.42s** — supports_predicted_class; evidence: unaligned vision feature rows 88–94; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: -0.5230 / -0.2394 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +1.9185 / +0.7809. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: -0.1013.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
He is the co-founder of Rossen and Vettese Limited and the former Executive Director of Uniform Final Examination (UFE) courses at Toronto's York University.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 15
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +1.708
|
||||
- Predicted-class confidence: 0.945
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/15.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.1609 | 0.579 | +0.7296 | 0.556 |
|
||||
| audio | +0.3684 | 0.184 | +0.1344 | 0.102 |
|
||||
| vision | +0.4769 | 0.238 | +0.4493 | 0.342 |
|
||||
|
||||
Shapley completeness residuals: class -4.44e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.2422 | -0.0633 |
|
||||
| Text + Vision | -0.3853 | -0.3292 |
|
||||
| Audio + Vision | -0.0929 | -0.0623 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 20–22 (relative progress 0.400–0.460), estimated clip interval 2.93–3.37s** — supports_predicted_class; evidence: prioriti
|
||||
- **text bins 41–42 (relative progress 0.820–0.860), estimated clip interval 6.01–6.31s** — supports_predicted_class; evidence: power to
|
||||
- **audio bins 3–5 (relative progress 0.060–0.120), estimated clip interval 0.44–0.88s** — supports_predicted_class; evidence: unaligned audio feature rows 8–17; review the linked source clip at the estimated relative span
|
||||
- **audio bins 11–11 (relative progress 0.220–0.240), estimated clip interval 1.61–1.76s** — supports_predicted_class; evidence: unaligned audio feature rows 31–34; review the linked source clip at the estimated relative span
|
||||
- **vision bins 14–16 (relative progress 0.280–0.340), estimated clip interval 2.05–2.49s** — supports_predicted_class; evidence: unaligned vision feature rows 30–37; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 49–49 (relative progress 0.980–1.000), estimated clip interval 7.19–7.33s** — supports_predicted_class; evidence: unaligned vision feature rows 107–109; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.2317 / +0.6047 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.0641 / +0.0459. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.6060.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
However, despite their poverty, the family prioritize education because they believed in its power to transform lives
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 16
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.797
|
||||
- Predicted-class confidence: 0.944
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/16.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +2.1145 | 0.855 | -1.7931 | 0.818 |
|
||||
| audio | +0.0127 | 0.005 | -0.0988 | 0.045 |
|
||||
| vision | +0.3469 | 0.140 | -0.2997 | 0.137 |
|
||||
|
||||
Shapley completeness residuals: class -4.44e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.1119 | +0.1550 |
|
||||
| Text + Vision | -0.2551 | +0.3214 |
|
||||
| Audio + Vision | -0.1040 | +0.1047 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 15–17 (relative progress 0.300–0.360), estimated clip interval 0.93–1.12s** — supports_predicted_class; evidence: s a
|
||||
- **text bins 34–35 (relative progress 0.680–0.720), estimated clip interval 2.12–2.24s** — supports_predicted_class; evidence: is a
|
||||
- **audio bins 32–34 (relative progress 0.640–0.700), estimated clip interval 1.99–2.18s** — opposes_predicted_class; evidence: unaligned audio feature rows 38–42; review the linked source clip at the estimated relative span
|
||||
- **audio bins 12–13 (relative progress 0.240–0.280), estimated clip interval 0.75–0.87s** — opposes_predicted_class; evidence: unaligned audio feature rows 14–16; review the linked source clip at the estimated relative span
|
||||
- **vision bins 9–11 (relative progress 0.180–0.240), estimated clip interval 0.56–0.75s** — supports_predicted_class; evidence: unaligned vision feature rows 7–10; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 47–48 (relative progress 0.940–0.980), estimated clip interval 2.93–3.05s** — supports_predicted_class; evidence: unaligned vision feature rows 41–43; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.7584 / +1.7979 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.0846 / -0.2086. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.3835.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
It's a terrible, this is a terrible movie
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 17
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +1.779
|
||||
- Predicted-class confidence: 0.951
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/17.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.7625 | 0.636 | +1.1566 | 0.628 |
|
||||
| audio | -0.3269 | 0.118 | -0.2282 | 0.124 |
|
||||
| vision | +0.6833 | 0.246 | +0.4559 | 0.248 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.2539 | +0.1653 |
|
||||
| Text + Vision | -0.6970 | -0.4395 |
|
||||
| Audio + Vision | +0.0401 | +0.0092 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 26–27 (relative progress 0.520–0.560), estimated clip interval 5.03–5.41s** — supports_predicted_class; evidence: and will
|
||||
- **text bins 2–2 (relative progress 0.040–0.060), estimated clip interval 0.39–0.58s** — supports_predicted_class; evidence: applying
|
||||
- **audio bins 32–33 (relative progress 0.640–0.680), estimated clip interval 6.19–6.57s** — opposes_predicted_class; evidence: unaligned audio feature rows 122–129; review the linked source clip at the estimated relative span
|
||||
- **audio bins 2–3 (relative progress 0.040–0.080), estimated clip interval 0.39–0.77s** — opposes_predicted_class; evidence: unaligned audio feature rows 7–15; review the linked source clip at the estimated relative span
|
||||
- **vision bins 46–49 (relative progress 0.920–1.000), estimated clip interval 8.89–9.67s** — supports_predicted_class; evidence: unaligned vision feature rows 131–142; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 29–29 (relative progress 0.580–0.600), estimated clip interval 5.61–5.80s** — supports_predicted_class; evidence: unaligned vision feature rows 82–85; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3403 / +1.1632 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.2587 / -0.0617. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.8566.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Applying these four design concepts to your presentations is simple, easy and will make people think you turned into a design guru.
|
||||
@@ -0,0 +1,57 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 18
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **neutral**
|
||||
- Intensity: -0.379
|
||||
- Predicted-class confidence: 0.507
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/18.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +0.5864 | 0.482 | -0.3991 | 0.487 |
|
||||
| audio | -0.2698 | 0.222 | +0.0225 | 0.028 |
|
||||
| vision | +0.3607 | 0.296 | -0.3975 | 0.485 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.0485 | -0.0700 |
|
||||
| Text + Vision | -0.1951 | +0.1509 |
|
||||
| Audio + Vision | -0.0033 | +0.2356 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 6–10 (relative progress 0.120–0.220), estimated clip interval 2.26–4.14s** — supports_predicted_class; evidence: - the first baltic cod fishery
|
||||
- **audio bins 13–14 (relative progress 0.260–0.300), estimated clip interval 4.89–5.64s** — opposes_predicted_class; evidence: unaligned audio feature rows 96–111; review the linked source clip at the estimated relative span
|
||||
- **audio bins 1–2 (relative progress 0.020–0.060), estimated clip interval 0.38–1.13s** — opposes_predicted_class; evidence: unaligned audio feature rows 7–22; review the linked source clip at the estimated relative span
|
||||
- **vision bins 18–20 (relative progress 0.360–0.420), estimated clip interval 6.77–7.90s** — supports_predicted_class; evidence: unaligned vision feature rows 100–117; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 40–41 (relative progress 0.800–0.840), estimated clip interval 15.05–15.80s** — supports_predicted_class; evidence: unaligned vision feature rows 224–235; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.2234 / +0.4210 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.2583 / -0.1544. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.3293.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
-And in Denmark - the first Baltic Cod fishery has been MSC – certified -Meanwhile, the Faeroese Mackerel Fishery has been denied MSC certification based on the fact that the fishery has failed to reach an agreement on mackerel quotas with Norway and the European Union.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 19
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.030
|
||||
- Predicted-class confidence: 0.496
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/19.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | -0.0481 | 0.042 | -0.4154 | 0.367 |
|
||||
| audio | -0.3937 | 0.346 | -0.3329 | 0.294 |
|
||||
| vision | +0.6959 | 0.612 | +0.3834 | 0.339 |
|
||||
|
||||
Shapley completeness residuals: class -5.55e-17, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.1825 | +0.1441 |
|
||||
| Text + Vision | -0.2654 | -0.0422 |
|
||||
| Audio + Vision | +0.0285 | +0.0474 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 35–36 (relative progress 0.700–0.740), estimated clip interval 9.58–10.13s** — opposes_predicted_class; evidence: over imperfections
|
||||
- **text bins 40–41 (relative progress 0.800–0.840), estimated clip interval 10.95–11.49s** — supports_predicted_class; evidence: but for the
|
||||
- **audio bins 16–18 (relative progress 0.320–0.380), estimated clip interval 4.38–5.20s** — opposes_predicted_class; evidence: unaligned audio feature rows 87–103; review the linked source clip at the estimated relative span
|
||||
- **audio bins 0–1 (relative progress 0.000–0.040), estimated clip interval 0.00–0.55s** — opposes_predicted_class; evidence: unaligned audio feature rows 0–10; review the linked source clip at the estimated relative span
|
||||
- **vision bins 31–33 (relative progress 0.620–0.680), estimated clip interval 8.48–9.30s** — supports_predicted_class; evidence: unaligned vision feature rows 126–138; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 27–28 (relative progress 0.540–0.580), estimated clip interval 7.39–7.94s** — supports_predicted_class; evidence: unaligned vision feature rows 110–118; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: -0.1179 / +0.0719 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.4099 / +0.0560. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.0836.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
People are surprisingly forgiving brands when they own up to mistakes, and unfortunately some haters out there love to point fingers and jump all over imperfections, but for the most part, people understand
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E0_EarlyConcat — 20
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +1.269
|
||||
- Predicted-class confidence: 0.865
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/20.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.7774 | 0.829 | +1.2266 | 0.738 |
|
||||
| audio | -0.3541 | 0.165 | -0.3939 | 0.237 |
|
||||
| vision | +0.0132 | 0.006 | +0.0415 | 0.025 |
|
||||
|
||||
Shapley completeness residuals: class -2.22e-16, intensity -1.11e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.3277 | +0.3908 |
|
||||
| Text + Vision | -0.0568 | +0.0263 |
|
||||
| Audio + Vision | +0.0454 | +0.0601 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 24–26 (relative progress 0.480–0.540), estimated clip interval 5.26–5.91s** — supports_predicted_class; evidence: ll have more live
|
||||
- **text bins 11–12 (relative progress 0.220–0.260), estimated clip interval 2.41–2.85s** — supports_predicted_class; evidence: of the description
|
||||
- **audio bins 33–34 (relative progress 0.660–0.700), estimated clip interval 7.23–7.67s** — opposes_predicted_class; evidence: unaligned audio feature rows 142–151; review the linked source clip at the estimated relative span
|
||||
- **audio bins 41–42 (relative progress 0.820–0.860), estimated clip interval 8.98–9.42s** — opposes_predicted_class; evidence: unaligned audio feature rows 177–185; review the linked source clip at the estimated relative span
|
||||
- **vision bins 42–44 (relative progress 0.840–0.900), estimated clip interval 9.20–9.86s** — supports_predicted_class; evidence: unaligned vision feature rows 136–145; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 15–16 (relative progress 0.300–0.340), estimated clip interval 3.29–3.72s** — opposes_predicted_class; evidence: unaligned vision feature rows 48–55; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.5846 / +1.8524 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.0170 / -0.1314. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.5152.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
And of course, click in the link of the description of this video for more, and we'll have more live updates and a stock market video (wrap-up) at the end of the day today.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 01
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **neutral**
|
||||
- Intensity: +0.251
|
||||
- Predicted-class confidence: 0.477
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/01.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.363 | 0.642 |
|
||||
| audio | 0.294 | 0.139 |
|
||||
| vision | 0.343 | 0.219 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 35–36 (relative progress 0.700–0.740), estimated clip interval 6.30–6.66s** — router_activity_not_signed_contribution; evidence: belt performs
|
||||
- **text bins 4–4 (relative progress 0.080–0.100), estimated clip interval 0.72–0.90s** — router_activity_not_signed_contribution; evidence: replacing these
|
||||
- **audio bins 32–32 (relative progress 0.640–0.660), estimated clip interval 5.76–5.94s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 114–118; review the linked source clip at the estimated relative span
|
||||
- **audio bins 35–35 (relative progress 0.700–0.720), estimated clip interval 6.30–6.48s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 125–128; review the linked source clip at the estimated relative span
|
||||
- **vision bins 4–5 (relative progress 0.080–0.120), estimated clip interval 0.72–1.08s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 10–16; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 35–36 (relative progress 0.700–0.740), estimated clip interval 6.30–6.66s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 94–99; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3703 / +1.0939 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.0722 / -0.1165. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.8309.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Replacing these wear components when replacing the timing belt is essential to ensuring the new belt performs to its mileage requirements
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 02
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **neutral**
|
||||
- Intensity: +0.110
|
||||
- Predicted-class confidence: 0.462
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/02.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.358 | 0.723 |
|
||||
| audio | 0.289 | 0.055 |
|
||||
| vision | 0.352 | 0.222 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 0–3 (relative progress 0.000–0.080), estimated clip interval 0.00–0.28s** — router_activity_not_signed_contribution; evidence: we
|
||||
- **text bins 48–48 (relative progress 0.960–0.980), estimated clip interval 3.35–3.42s** — router_activity_not_signed_contribution; evidence: We want to live by each other’s happiness - not by each other’s misery.
|
||||
- **audio bins 19–20 (relative progress 0.380–0.420), estimated clip interval 1.33–1.47s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 25–27; review the linked source clip at the estimated relative span
|
||||
- **audio bins 9–10 (relative progress 0.180–0.220), estimated clip interval 0.63–0.77s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 11–14; review the linked source clip at the estimated relative span
|
||||
- **vision bins 48–49 (relative progress 0.960–1.000), estimated clip interval 3.35–3.49s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 48–49; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 1–2 (relative progress 0.020–0.060), estimated clip interval 0.07–0.21s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 1–2; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.2126 / +0.9516 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.1482 / +0.1970. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.6623.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
We want to live by each other’s happiness - not by each other’s misery.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 03
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.610
|
||||
- Predicted-class confidence: 0.688
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/03.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.359 | 0.723 |
|
||||
| audio | 0.292 | 0.011 |
|
||||
| vision | 0.349 | 0.266 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 33–34 (relative progress 0.660–0.700), estimated clip interval 6.38–6.76s** — router_activity_not_signed_contribution; evidence: income into
|
||||
- **text bins 9–9 (relative progress 0.180–0.200), estimated clip interval 1.74–1.93s** — router_activity_not_signed_contribution; evidence: ##er
|
||||
- **audio bins 31–34 (relative progress 0.620–0.700), estimated clip interval 5.99–6.76s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 118–133; review the linked source clip at the estimated relative span
|
||||
- **audio bins 0–0 (relative progress 0.000–0.020), estimated clip interval 0.00–0.19s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 0–3; review the linked source clip at the estimated relative span
|
||||
- **vision bins 37–39 (relative progress 0.740–0.800), estimated clip interval 7.15–7.73s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 106–115; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 48–48 (relative progress 0.960–0.980), estimated clip interval 9.28–9.47s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 138–141; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3234 / +1.2305 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.1027 / +0.1196. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.8736.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
There's one lender at the moment which I think is just Bankwest who don't take rental income into account, they take rental yield into account.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 04
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.942
|
||||
- Predicted-class confidence: 0.784
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/04.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.359 | 0.789 |
|
||||
| audio | 0.291 | 0.196 |
|
||||
| vision | 0.349 | 0.015 |
|
||||
|
||||
Router–Shapley Spearman: 0.5; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 44–47 (relative progress 0.880–0.960), estimated clip interval 6.15–6.71s** — router_activity_not_signed_contribution; evidence: absolutely not.
|
||||
- **text bins 21–21 (relative progress 0.420–0.440), estimated clip interval 2.94–3.08s** — router_activity_not_signed_contribution; evidence: i kiss
|
||||
- **audio bins 13–15 (relative progress 0.260–0.320), estimated clip interval 1.82–2.24s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 35–43; review the linked source clip at the estimated relative span
|
||||
- **audio bins 46–46 (relative progress 0.920–0.940), estimated clip interval 6.43–6.57s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 126–128; review the linked source clip at the estimated relative span
|
||||
- **vision bins 39–42 (relative progress 0.780–0.860), estimated clip interval 5.46–6.01s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 80–88; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 7–7 (relative progress 0.140–0.160), estimated clip interval 0.98–1.12s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 14–16; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3890 / +1.5730 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.4015 / +0.1569. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.4743.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
If I blow it at the team exercise, should I kiss my chances of cheering "GO BLUE" goodbye?] Absolutely not.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 05
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.871
|
||||
- Predicted-class confidence: 0.837
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/05.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.359 | 0.456 |
|
||||
| audio | 0.291 | 0.111 |
|
||||
| vision | 0.350 | 0.433 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 43–45 (relative progress 0.860–0.920), estimated clip interval 5.25–5.61s** — router_activity_not_signed_contribution; evidence: marketing.
|
||||
- **text bins 0–0 (relative progress 0.000–0.020), estimated clip interval 0.00–0.12s** — router_activity_not_signed_contribution; evidence: Hi, my name is Chloe, video marketer for Red Wagon Marketing.
|
||||
- **audio bins 43–46 (relative progress 0.860–0.940), estimated clip interval 5.25–5.73s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 103–112; review the linked source clip at the estimated relative span
|
||||
- **audio bins 23–23 (relative progress 0.460–0.480), estimated clip interval 2.81–2.93s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 55–57; review the linked source clip at the estimated relative span
|
||||
- **vision bins 3–5 (relative progress 0.060–0.120), estimated clip interval 0.37–0.73s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 5–10; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 17–17 (relative progress 0.340–0.360), estimated clip interval 2.07–2.20s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 30–32; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.0746 / +0.1400 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.9288 / +0.5481. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.4152.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Hi, my name is Chloe, video marketer for Red Wagon Marketing.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 06
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.790
|
||||
- Predicted-class confidence: 0.805
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/06.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.359 | 0.801 |
|
||||
| audio | 0.290 | 0.121 |
|
||||
| vision | 0.351 | 0.078 |
|
||||
|
||||
Router–Shapley Spearman: 0.5; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 40–40 (relative progress 0.800–0.820), estimated clip interval 6.64–6.81s** — router_activity_not_signed_contribution; evidence: check out
|
||||
- **text bins 0–0 (relative progress 0.000–0.020), estimated clip interval 0.00–0.17s** — router_activity_not_signed_contribution; evidence: If you're a fan of dancing in that sense, just like to watch people dance, see impressive dance moves then you might want to check out this movie solely for that
|
||||
- **audio bins 43–44 (relative progress 0.860–0.900), estimated clip interval 7.14–7.47s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 140–146; review the linked source clip at the estimated relative span
|
||||
- **audio bins 39–40 (relative progress 0.780–0.820), estimated clip interval 6.48–6.81s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 127–133; review the linked source clip at the estimated relative span
|
||||
- **vision bins 46–49 (relative progress 0.920–1.000), estimated clip interval 7.64–8.31s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 113–122; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 13–13 (relative progress 0.260–0.280), estimated clip interval 2.16–2.33s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 31–34; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.5262 / +1.6266 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.4002 / -0.1076. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.3422.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
If you're a fan of dancing in that sense, just like to watch people dance, see impressive dance moves then you might want to check out this movie solely for that
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 07
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.580
|
||||
- Predicted-class confidence: 0.734
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/07.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.362 | 0.838 |
|
||||
| audio | 0.294 | 0.153 |
|
||||
| vision | 0.344 | 0.009 |
|
||||
|
||||
Router–Shapley Spearman: 0.5; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 31–32 (relative progress 0.620–0.660), estimated clip interval 13.24–14.09s** — router_activity_not_signed_contribution; evidence: future of
|
||||
- **text bins 0–0 (relative progress 0.000–0.020), estimated clip interval 0.00–0.43s** — router_activity_not_signed_contribution; evidence: As Linn’s associate editor Michael Baadke reports in our November 28 issue, attendees “enthusiastically discussed the topics of growing the hobby, the future of stamp shows, and dealers and philatelic partnerships, along with ways the leading organizations involved in the stamp hobby can work together to make it succeed and grow
|
||||
- **audio bins 28–29 (relative progress 0.560–0.600), estimated clip interval 11.95–12.81s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 238–254; review the linked source clip at the estimated relative span
|
||||
- **audio bins 34–35 (relative progress 0.680–0.720), estimated clip interval 14.52–15.37s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 289–305; review the linked source clip at the estimated relative span
|
||||
- **vision bins 18–19 (relative progress 0.360–0.400), estimated clip interval 7.69–8.54s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 114–127; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 36–37 (relative progress 0.720–0.760), estimated clip interval 15.37–16.22s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 229–242; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.1449 / +0.7703 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.4447 / +0.0732. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.6005.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
As Linn’s associate editor Michael Baadke reports in our November 28 issue, attendees “enthusiastically discussed the topics of growing the hobby, the future of stamp shows, and dealers and philatelic partnerships, along with ways the leading organizations involved in the stamp hobby can work together to make it succeed and grow
|
||||
@@ -0,0 +1,49 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 08
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.712
|
||||
- Predicted-class confidence: 0.812
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/08.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.363 | 0.888 |
|
||||
| audio | 0.294 | 0.090 |
|
||||
| vision | 0.343 | 0.022 |
|
||||
|
||||
Router–Shapley Spearman: 0.5; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 42–43 (relative progress 0.840–0.880), estimated clip interval 4.27–4.48s** — router_activity_not_signed_contribution; evidence: grand challenge
|
||||
- **text bins 47–48 (relative progress 0.940–0.980), estimated clip interval 4.78–4.98s** — router_activity_not_signed_contribution; evidence: That brings us to tonight, the Universal Design Grand Challenge
|
||||
- **audio bins 23–27 (relative progress 0.460–0.560), estimated clip interval 2.34–2.85s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 45–54; review the linked source clip at the estimated relative span
|
||||
- **vision bins 47–49 (relative progress 0.940–1.000), estimated clip interval 4.78–5.09s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 69–73; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 42–43 (relative progress 0.840–0.880), estimated clip interval 4.27–4.48s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 62–65; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.0763 / +1.0180 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.8223 / +0.0351. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.8006.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
That brings us to tonight, the Universal Design Grand Challenge
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 09
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.848
|
||||
- Predicted-class confidence: 0.918
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/09.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.359 | 0.940 |
|
||||
| audio | 0.291 | 0.052 |
|
||||
| vision | 0.350 | 0.008 |
|
||||
|
||||
Router–Shapley Spearman: 0.5; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 8–10 (relative progress 0.160–0.220), estimated clip interval 0.81–1.11s** — router_activity_not_signed_contribution; evidence: ##h )
|
||||
- **text bins 4–4 (relative progress 0.080–0.100), estimated clip interval 0.40–0.50s** — router_activity_not_signed_contribution; evidence: (
|
||||
- **audio bins 8–10 (relative progress 0.160–0.220), estimated clip interval 0.81–1.11s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 15–21; review the linked source clip at the estimated relative span
|
||||
- **audio bins 2–3 (relative progress 0.040–0.080), estimated clip interval 0.20–0.40s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 3–7; review the linked source clip at the estimated relative span
|
||||
- **vision bins 25–27 (relative progress 0.500–0.560), estimated clip interval 2.52–2.82s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 37–41; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 12–12 (relative progress 0.240–0.260), estimated clip interval 1.21–1.31s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 18–19; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3663 / +1.5409 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.4417 / +0.2561. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.3679.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
(uhh) I did not like this movie at all, I would not recommend it
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 10
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.260
|
||||
- Predicted-class confidence: 0.862
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/10.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.359 | 0.665 |
|
||||
| audio | 0.290 | 0.213 |
|
||||
| vision | 0.351 | 0.123 |
|
||||
|
||||
Router–Shapley Spearman: 0.5; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 5–7 (relative progress 0.100–0.160), estimated clip interval 0.99–1.59s** — router_activity_not_signed_contribution; evidence: ) and you
|
||||
- **text bins 48–48 (relative progress 0.960–0.980), estimated clip interval 9.54–9.74s** — router_activity_not_signed_contribution; evidence: terrible
|
||||
- **audio bins 5–6 (relative progress 0.100–0.140), estimated clip interval 0.99–1.39s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 19–27; review the linked source clip at the estimated relative span
|
||||
- **audio bins 28–29 (relative progress 0.560–0.600), estimated clip interval 5.57–5.96s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 109–116; review the linked source clip at the estimated relative span
|
||||
- **vision bins 10–11 (relative progress 0.200–0.240), estimated clip interval 1.99–2.39s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 29–35; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 23–23 (relative progress 0.460–0.480), estimated clip interval 4.57–4.77s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 67–70; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3188 / +1.7735 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.7996 / +0.2699. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.7536.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
(umm) And you know I really do like to see fluffy chick flicks sometimes so I'm not against that but this one was pretty terrible
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 11
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.492
|
||||
- Predicted-class confidence: 0.577
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/11.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.360 | 0.543 |
|
||||
| audio | 0.288 | 0.133 |
|
||||
| vision | 0.352 | 0.325 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 31–34 (relative progress 0.620–0.700), estimated clip interval 3.60–4.06s** — router_activity_not_signed_contribution; evidence: film if i
|
||||
- **text bins 40–40 (relative progress 0.800–0.820), estimated clip interval 4.64–4.76s** — router_activity_not_signed_contribution; evidence: was a
|
||||
- **audio bins 1–2 (relative progress 0.020–0.060), estimated clip interval 0.12–0.35s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 2–6; review the linked source clip at the estimated relative span
|
||||
- **audio bins 5–5 (relative progress 0.100–0.120), estimated clip interval 0.58–0.70s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 11–13; review the linked source clip at the estimated relative span
|
||||
- **vision bins 32–33 (relative progress 0.640–0.680), estimated clip interval 3.71–3.95s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 55–58; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 7–7 (relative progress 0.140–0.160), estimated clip interval 0.81–0.93s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 12–13; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: -0.1744 / +0.4081 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.8332 / -0.1700. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.3823.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
I would be ashamed to have made this film if I was a director
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 12
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.088
|
||||
- Predicted-class confidence: 0.814
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/12.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.363 | 0.518 |
|
||||
| audio | 0.291 | 0.198 |
|
||||
| vision | 0.346 | 0.284 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 31–33 (relative progress 0.620–0.680), estimated clip interval 8.07–8.85s** — router_activity_not_signed_contribution; evidence: , they have
|
||||
- **text bins 40–41 (relative progress 0.800–0.840), estimated clip interval 10.41–10.93s** — router_activity_not_signed_contribution; evidence: , and
|
||||
- **audio bins 10–11 (relative progress 0.200–0.240), estimated clip interval 2.60–3.12s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 51–61; review the linked source clip at the estimated relative span
|
||||
- **audio bins 28–28 (relative progress 0.560–0.580), estimated clip interval 7.29–7.55s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 143–149; review the linked source clip at the estimated relative span
|
||||
- **vision bins 31–33 (relative progress 0.620–0.680), estimated clip interval 8.07–8.85s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 119–131; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 40–41 (relative progress 0.800–0.840), estimated clip interval 10.41–10.93s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 154–162; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.1360 / +0.8681 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.3464 / +0.3088. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.0633.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Or worse, an individual previously had good credit, but usually by no fault of their own, or perhaps by fault of their own, they have let their credit sag, and credit scores is very low.
|
||||
@@ -0,0 +1,49 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 13
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.421
|
||||
- Predicted-class confidence: 0.525
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/13.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.363 | 0.515 |
|
||||
| audio | 0.296 | 0.126 |
|
||||
| vision | 0.340 | 0.359 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 3–4 (relative progress 0.060–0.100), estimated clip interval 0.52–0.87s** — router_activity_not_signed_contribution; evidence: for example,
|
||||
- **text bins 14–15 (relative progress 0.280–0.320), estimated clip interval 2.43–2.78s** — router_activity_not_signed_contribution; evidence: set of data
|
||||
- **audio bins 1–5 (relative progress 0.020–0.120), estimated clip interval 0.17–1.04s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 3–20; review the linked source clip at the estimated relative span
|
||||
- **vision bins 14–15 (relative progress 0.280–0.320), estimated clip interval 2.43–2.78s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 5–5; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 20–20 (relative progress 0.400–0.420), estimated clip interval 3.47–3.65s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 7–7; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.0939 / +0.2419 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.0028 / -0.1247. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.2177.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
For example, I could take a set of data and from that data, I can find a relationship between any two of the given factors or more.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 14
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.526
|
||||
- Predicted-class confidence: 0.582
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/14.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.360 | 0.115 |
|
||||
| audio | 0.292 | 0.177 |
|
||||
| vision | 0.348 | 0.708 |
|
||||
|
||||
Router–Shapley Spearman: -0.5; top modality agreement: False.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 0–2 (relative progress 0.000–0.060), estimated clip interval 0.00–0.62s** — router_activity_not_signed_contribution; evidence: he is
|
||||
- **text bins 46–46 (relative progress 0.920–0.940), estimated clip interval 9.52–9.73s** — router_activity_not_signed_contribution; evidence: university
|
||||
- **audio bins 0–2 (relative progress 0.000–0.060), estimated clip interval 0.00–0.62s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 0–12; review the linked source clip at the estimated relative span
|
||||
- **audio bins 46–46 (relative progress 0.920–0.940), estimated clip interval 9.52–9.73s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 187–191; review the linked source clip at the estimated relative span
|
||||
- **vision bins 5–7 (relative progress 0.100–0.160), estimated clip interval 1.04–1.66s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 15–24; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 32–32 (relative progress 0.640–0.660), estimated clip interval 6.62–6.83s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 97–100; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.0264 / -0.3809 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.4573 / +0.3264. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: -0.2121.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
He is the co-founder of Rossen and Vettese Limited and the former Executive Director of Uniform Final Examination (UFE) courses at Toronto's York University.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 15
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +1.053
|
||||
- Predicted-class confidence: 0.893
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/15.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.359 | 0.496 |
|
||||
| audio | 0.292 | 0.244 |
|
||||
| vision | 0.349 | 0.260 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 15–16 (relative progress 0.300–0.340), estimated clip interval 2.20–2.49s** — router_activity_not_signed_contribution; evidence: , the
|
||||
- **text bins 27–28 (relative progress 0.540–0.580), estimated clip interval 3.96–4.25s** — router_activity_not_signed_contribution; evidence: education because
|
||||
- **audio bins 8–9 (relative progress 0.160–0.200), estimated clip interval 1.17–1.47s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 23–28; review the linked source clip at the estimated relative span
|
||||
- **audio bins 15–15 (relative progress 0.300–0.320), estimated clip interval 2.20–2.35s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 43–46; review the linked source clip at the estimated relative span
|
||||
- **vision bins 47–48 (relative progress 0.940–0.980), estimated clip interval 6.89–7.19s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 103–107; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 23–24 (relative progress 0.460–0.500), estimated clip interval 3.37–3.67s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 50–54; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.1767 / +0.5629 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.9233 / +0.5012. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.4848.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
However, despite their poverty, the family prioritize education because they believed in its power to transform lives
|
||||
@@ -0,0 +1,49 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 16
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.836
|
||||
- Predicted-class confidence: 0.914
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/16.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.370 | 0.876 |
|
||||
| audio | 0.304 | 0.050 |
|
||||
| vision | 0.326 | 0.075 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 2–4 (relative progress 0.040–0.100), estimated clip interval 0.12–0.31s** — router_activity_not_signed_contribution; evidence: it
|
||||
- **text bins 46–46 (relative progress 0.920–0.940), estimated clip interval 2.86–2.93s** — router_activity_not_signed_contribution; evidence: movie
|
||||
- **audio bins 1–5 (relative progress 0.020–0.120), estimated clip interval 0.06–0.37s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 1–7; review the linked source clip at the estimated relative span
|
||||
- **vision bins 27–30 (relative progress 0.540–0.620), estimated clip interval 1.68–1.93s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 23–27; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 15–15 (relative progress 0.300–0.320), estimated clip interval 0.93–1.00s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 13–14; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.2071 / +1.2711 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.5664 / +0.1907. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.1494.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
It's a terrible, this is a terrible movie
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 17
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.928
|
||||
- Predicted-class confidence: 0.884
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/17.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.360 | 0.721 |
|
||||
| audio | 0.289 | 0.079 |
|
||||
| vision | 0.351 | 0.200 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 10–11 (relative progress 0.200–0.240), estimated clip interval 1.93–2.32s** — router_activity_not_signed_contribution; evidence: concepts to
|
||||
- **text bins 48–49 (relative progress 0.960–1.000), estimated clip interval 9.28–9.67s** — router_activity_not_signed_contribution; evidence: .
|
||||
- **audio bins 37–38 (relative progress 0.740–0.780), estimated clip interval 7.15–7.54s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 141–148; review the linked source clip at the estimated relative span
|
||||
- **audio bins 2–2 (relative progress 0.040–0.060), estimated clip interval 0.39–0.58s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 7–11; review the linked source clip at the estimated relative span
|
||||
- **vision bins 46–47 (relative progress 0.920–0.960), estimated clip interval 8.89–9.28s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 131–137; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 17–18 (relative progress 0.340–0.380), estimated clip interval 3.29–3.67s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 48–54; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3064 / +1.3990 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.5635 / -0.0240. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.2303.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Applying these four design concepts to your presentations is simple, easy and will make people think you turned into a design guru.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 18
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.171
|
||||
- Predicted-class confidence: 0.490
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/18.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.359 | 0.380 |
|
||||
| audio | 0.291 | 0.178 |
|
||||
| vision | 0.350 | 0.441 |
|
||||
|
||||
Router–Shapley Spearman: 0.5; top modality agreement: False.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 38–38 (relative progress 0.760–0.780), estimated clip interval 14.29–14.67s** — router_activity_not_signed_contribution; evidence: fishery
|
||||
- **text bins 7–7 (relative progress 0.140–0.160), estimated clip interval 2.63–3.01s** — router_activity_not_signed_contribution; evidence: first
|
||||
- **audio bins 38–38 (relative progress 0.760–0.780), estimated clip interval 14.29–14.67s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 282–290; review the linked source clip at the estimated relative span
|
||||
- **audio bins 4–4 (relative progress 0.080–0.100), estimated clip interval 1.50–1.88s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 29–37; review the linked source clip at the estimated relative span
|
||||
- **vision bins 10–11 (relative progress 0.200–0.240), estimated clip interval 3.76–4.51s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 56–67; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 2–2 (relative progress 0.040–0.060), estimated clip interval 0.75–1.13s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 11–16; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: -0.1406 / -0.3074 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +1.0792 / +0.8558. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: -0.0964.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
-And in Denmark - the first Baltic Cod fishery has been MSC – certified -Meanwhile, the Faeroese Mackerel Fishery has been denied MSC certification based on the fact that the fishery has failed to reach an agreement on mackerel quotas with Norway and the European Union.
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 19
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.047
|
||||
- Predicted-class confidence: 0.474
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/19.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.358 | 0.080 |
|
||||
| audio | 0.291 | 0.053 |
|
||||
| vision | 0.351 | 0.867 |
|
||||
|
||||
Router–Shapley Spearman: 0.5; top modality agreement: False.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 45–46 (relative progress 0.900–0.940), estimated clip interval 12.31–12.86s** — router_activity_not_signed_contribution; evidence: part, people
|
||||
- **text bins 30–30 (relative progress 0.600–0.620), estimated clip interval 8.21–8.48s** — router_activity_not_signed_contribution; evidence: fingers and
|
||||
- **audio bins 45–46 (relative progress 0.900–0.940), estimated clip interval 12.31–12.86s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 244–255; review the linked source clip at the estimated relative span
|
||||
- **audio bins 30–31 (relative progress 0.600–0.640), estimated clip interval 8.21–8.76s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 163–174; review the linked source clip at the estimated relative span
|
||||
- **vision bins 7–9 (relative progress 0.140–0.200), estimated clip interval 1.92–2.74s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 28–40; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 20–20 (relative progress 0.400–0.420), estimated clip interval 5.47–5.75s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 81–85; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.2485 / -0.0569 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.0621 / +0.5542. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.1364.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
People are surprisingly forgiving brands when they own up to mistakes, and unfortunately some haters out there love to point fingers and jump all over imperfections, but for the most part, people understand
|
||||
@@ -0,0 +1,50 @@
|
||||
# Q3 explanation card — E1_MoFE_Router — 20
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.760
|
||||
- Predicted-class confidence: 0.835
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/20.mp4
|
||||
|
||||
## Router profile (intrinsic routing signal, not prediction contribution)
|
||||
|
||||
| Modality | Router exposure share | Exact Shapley absolute share |
|
||||
|---|---:|---:|
|
||||
| text | 0.359 | 0.936 |
|
||||
| audio | 0.291 | 0.025 |
|
||||
| vision | 0.350 | 0.039 |
|
||||
|
||||
Router–Shapley Spearman: 1.0; top modality agreement: True.
|
||||
Router values describe mixture routing. The counterfactual scores below test whether that routing signal tracks model behavior.
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 15–16 (relative progress 0.300–0.340), estimated clip interval 3.29–3.72s** — router_activity_not_signed_contribution; evidence: this video for
|
||||
- **text bins 33–34 (relative progress 0.660–0.700), estimated clip interval 7.23–7.67s** — router_activity_not_signed_contribution; evidence: market video (
|
||||
- **audio bins 33–35 (relative progress 0.660–0.720), estimated clip interval 7.23–7.89s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 142–155; review the linked source clip at the estimated relative span
|
||||
- **audio bins 16–16 (relative progress 0.320–0.340), estimated clip interval 3.50–3.72s** — router_activity_not_signed_contribution; evidence: unaligned audio feature rows 69–73; review the linked source clip at the estimated relative span
|
||||
- **vision bins 7–9 (relative progress 0.140–0.200), estimated clip interval 1.53–2.19s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 22–32; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 5–5 (relative progress 0.100–0.120), estimated clip interval 1.10–1.31s** — router_activity_not_signed_contribution; evidence: unaligned vision feature rows 16–19; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.5186 / +1.6799 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.5129 / +0.0639. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.3739.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
And of course, click in the link of the description of this video for more, and we'll have more live updates and a stock market video (wrap-up) at the end of the day today.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 01
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **neutral**
|
||||
- Intensity: +0.251
|
||||
- Predicted-class confidence: 0.477
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/01.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +0.9563 | 0.642 | +0.6618 | 0.603 |
|
||||
| audio | -0.2072 | 0.139 | +0.1829 | 0.167 |
|
||||
| vision | -0.3270 | 0.219 | -0.2536 | 0.231 |
|
||||
|
||||
Shapley completeness residuals: class -5.55e-17, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.1493 | -0.1375 |
|
||||
| Text + Vision | +0.1915 | +0.2039 |
|
||||
| Audio + Vision | +0.0868 | +0.0537 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 15–18 (relative progress 0.300–0.380), estimated clip interval 2.70–3.42s** — supports_predicted_class; evidence: the timing belt
|
||||
- **text bins 32–32 (relative progress 0.640–0.660), estimated clip interval 5.76–5.94s** — supports_predicted_class; evidence: new
|
||||
- **audio bins 1–2 (relative progress 0.020–0.060), estimated clip interval 0.18–0.54s** — opposes_predicted_class; evidence: unaligned audio feature rows 3–10; review the linked source clip at the estimated relative span
|
||||
- **audio bins 16–17 (relative progress 0.320–0.360), estimated clip interval 2.88–3.24s** — opposes_predicted_class; evidence: unaligned audio feature rows 57–64; review the linked source clip at the estimated relative span
|
||||
- **vision bins 5–8 (relative progress 0.100–0.180), estimated clip interval 0.90–1.62s** — opposes_predicted_class; evidence: unaligned vision feature rows 13–24; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 46–46 (relative progress 0.920–0.940), estimated clip interval 8.28–8.46s** — opposes_predicted_class; evidence: unaligned vision feature rows 124–126; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.4212 / +1.1640 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.0226 / -0.0818. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.8752.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Replacing these wear components when replacing the timing belt is essential to ensuring the new belt performs to its mileage requirements
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 02
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **neutral**
|
||||
- Intensity: +0.110
|
||||
- Predicted-class confidence: 0.462
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/02.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +0.6762 | 0.723 | +0.8219 | 0.671 |
|
||||
| audio | -0.0516 | 0.055 | +0.0157 | 0.013 |
|
||||
| vision | -0.2072 | 0.222 | -0.3872 | 0.316 |
|
||||
|
||||
Shapley completeness residuals: class -5.55e-17, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.1744 | +0.0061 |
|
||||
| Text + Vision | +0.2674 | +0.1294 |
|
||||
| Audio + Vision | +0.0312 | -0.0598 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 26–28 (relative progress 0.520–0.580), estimated clip interval 1.82–2.03s** — supports_predicted_class; evidence: happiness - not
|
||||
- **text bins 39–39 (relative progress 0.780–0.800), estimated clip interval 2.72–2.79s** — supports_predicted_class; evidence: ’
|
||||
- **audio bins 31–33 (relative progress 0.620–0.680), estimated clip interval 2.17–2.38s** — supports_predicted_class; evidence: unaligned audio feature rows 40–44; review the linked source clip at the estimated relative span
|
||||
- **audio bins 35–36 (relative progress 0.700–0.740), estimated clip interval 2.45–2.58s** — supports_predicted_class; evidence: unaligned audio feature rows 46–48; review the linked source clip at the estimated relative span
|
||||
- **vision bins 40–43 (relative progress 0.800–0.880), estimated clip interval 2.79–3.07s** — opposes_predicted_class; evidence: unaligned vision feature rows 40–43; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 17–17 (relative progress 0.340–0.360), estimated clip interval 1.19–1.26s** — supports_predicted_class; evidence: unaligned vision feature rows 17–18; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3460 / +1.0726 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.3472 / +0.2407. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.7905.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
We want to live by each other’s happiness - not by each other’s misery.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 03
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.610
|
||||
- Predicted-class confidence: 0.688
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/03.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.2257 | 0.723 | -0.6208 | 0.639 |
|
||||
| audio | +0.0185 | 0.011 | +0.0243 | 0.025 |
|
||||
| vision | -0.4508 | 0.266 | +0.3268 | 0.336 |
|
||||
|
||||
Shapley completeness residuals: class -1.11e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.1065 | +0.1757 |
|
||||
| Text + Vision | +0.5189 | -0.3861 |
|
||||
| Audio + Vision | +0.1038 | -0.1471 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 44–45 (relative progress 0.880–0.920), estimated clip interval 8.50–8.89s** — supports_predicted_class; evidence: yield into account
|
||||
- **text bins 35–36 (relative progress 0.700–0.740), estimated clip interval 6.76–7.15s** — supports_predicted_class; evidence: into account
|
||||
- **audio bins 7–10 (relative progress 0.140–0.220), estimated clip interval 1.35–2.13s** — opposes_predicted_class; evidence: unaligned audio feature rows 26–42; review the linked source clip at the estimated relative span
|
||||
- **audio bins 3–3 (relative progress 0.060–0.080), estimated clip interval 0.58–0.77s** — opposes_predicted_class; evidence: unaligned audio feature rows 11–15; review the linked source clip at the estimated relative span
|
||||
- **vision bins 42–44 (relative progress 0.840–0.900), estimated clip interval 8.12–8.70s** — supports_predicted_class; evidence: unaligned vision feature rows 120–129; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 7–8 (relative progress 0.140–0.180), estimated clip interval 1.35–1.74s** — supports_predicted_class; evidence: unaligned vision feature rows 20–25; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.7735 / +1.4775 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.6596 / -0.0313. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.1804.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
There's one lender at the moment which I think is just Bankwest who don't take rental income into account, they take rental yield into account.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 04
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.942
|
||||
- Predicted-class confidence: 0.784
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/04.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.6040 | 0.789 | -1.1000 | 0.688 |
|
||||
| audio | -0.3985 | 0.196 | +0.4059 | 0.254 |
|
||||
| vision | -0.0300 | 0.015 | +0.0925 | 0.058 |
|
||||
|
||||
Shapley completeness residuals: class -2.22e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.1932 | -0.1284 |
|
||||
| Text + Vision | +0.2459 | -0.1719 |
|
||||
| Audio + Vision | +0.2132 | -0.2422 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 43–45 (relative progress 0.860–0.920), estimated clip interval 6.01–6.43s** — supports_predicted_class; evidence: absolutely not
|
||||
- **text bins 6–7 (relative progress 0.120–0.160), estimated clip interval 0.84–1.12s** — supports_predicted_class; evidence: blow it
|
||||
- **audio bins 39–42 (relative progress 0.780–0.860), estimated clip interval 5.46–6.01s** — opposes_predicted_class; evidence: unaligned audio feature rows 106–117; review the linked source clip at the estimated relative span
|
||||
- **audio bins 12–12 (relative progress 0.240–0.260), estimated clip interval 1.68–1.82s** — supports_predicted_class; evidence: unaligned audio feature rows 32–35; review the linked source clip at the estimated relative span
|
||||
- **vision bins 39–41 (relative progress 0.780–0.840), estimated clip interval 5.46–5.87s** — supports_predicted_class; evidence: unaligned vision feature rows 80–86; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 31–32 (relative progress 0.620–0.660), estimated clip interval 4.34–4.62s** — supports_predicted_class; evidence: unaligned vision feature rows 63–67; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.7235 / +2.1628 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.3664 / -0.2742. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.6747.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
If I blow it at the team exercise, should I kiss my chances of cheering "GO BLUE" goodbye?] Absolutely not.
|
||||
@@ -0,0 +1,54 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 05
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.871
|
||||
- Predicted-class confidence: 0.837
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/05.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +0.8392 | 0.456 | +0.6213 | 0.513 |
|
||||
| audio | +0.2036 | 0.111 | +0.0360 | 0.030 |
|
||||
| vision | +0.7976 | 0.433 | +0.5540 | 0.457 |
|
||||
|
||||
Shapley completeness residuals: class -2.22e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.1651 | -0.0478 |
|
||||
| Text + Vision | -0.4251 | -0.3215 |
|
||||
| Audio + Vision | -0.1776 | -0.0663 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 6–10 (relative progress 0.120–0.220), estimated clip interval 0.73–1.34s** — supports_predicted_class; evidence: , my
|
||||
- **audio bins 1–5 (relative progress 0.020–0.120), estimated clip interval 0.12–0.73s** — opposes_predicted_class; evidence: unaligned audio feature rows 2–14; review the linked source clip at the estimated relative span
|
||||
- **vision bins 43–47 (relative progress 0.860–0.960), estimated clip interval 5.25–5.86s** — supports_predicted_class; evidence: unaligned vision feature rows 77–86; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3077 / +0.3107 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.2027 / +0.2257. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.4245.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Hi, my name is Chloe, video marketer for Red Wagon Marketing.
|
||||
@@ -0,0 +1,56 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 06
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.790
|
||||
- Predicted-class confidence: 0.805
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/06.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.8127 | 0.801 | +1.1609 | 0.800 |
|
||||
| audio | -0.2741 | 0.121 | -0.1604 | 0.111 |
|
||||
| vision | +0.1757 | 0.078 | +0.1295 | 0.089 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity -2.22e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.0391 | +0.1084 |
|
||||
| Text + Vision | -0.1121 | -0.1039 |
|
||||
| Audio + Vision | -0.1589 | -0.2370 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 21–23 (relative progress 0.420–0.480), estimated clip interval 3.49–3.99s** — supports_predicted_class; evidence: to watch people
|
||||
- **text bins 7–8 (relative progress 0.140–0.180), estimated clip interval 1.16–1.49s** — supports_predicted_class; evidence: a fan
|
||||
- **audio bins 42–44 (relative progress 0.840–0.900), estimated clip interval 6.98–7.47s** — opposes_predicted_class; evidence: unaligned audio feature rows 136–146; review the linked source clip at the estimated relative span
|
||||
- **audio bins 14–15 (relative progress 0.280–0.320), estimated clip interval 2.33–2.66s** — opposes_predicted_class; evidence: unaligned audio feature rows 45–52; review the linked source clip at the estimated relative span
|
||||
- **vision bins 40–44 (relative progress 0.800–0.900), estimated clip interval 6.64–7.47s** — opposes_predicted_class; evidence: unaligned vision feature rows 98–110; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.6819 / +1.6836 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.0933 / -0.2532. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.2527.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
If you're a fan of dancing in that sense, just like to watch people dance, see impressive dance moves then you might want to check out this movie solely for that
|
||||
@@ -0,0 +1,57 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 07
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.580
|
||||
- Predicted-class confidence: 0.734
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/07.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.2243 | 0.838 | +1.0975 | 0.705 |
|
||||
| audio | +0.2240 | 0.153 | +0.1413 | 0.091 |
|
||||
| vision | +0.0135 | 0.009 | -0.3190 | 0.205 |
|
||||
|
||||
Shapley completeness residuals: class -2.22e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.2760 | -0.1914 |
|
||||
| Text + Vision | +0.0270 | +0.2232 |
|
||||
| Audio + Vision | -0.1128 | -0.0191 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 22–26 (relative progress 0.440–0.540), estimated clip interval 9.39–11.53s** — supports_predicted_class; evidence: discussed the topics of growing
|
||||
- **audio bins 25–26 (relative progress 0.500–0.540), estimated clip interval 10.67–11.53s** — opposes_predicted_class; evidence: unaligned audio feature rows 212–229; review the linked source clip at the estimated relative span
|
||||
- **audio bins 30–31 (relative progress 0.600–0.640), estimated clip interval 12.81–13.66s** — opposes_predicted_class; evidence: unaligned audio feature rows 255–271; review the linked source clip at the estimated relative span
|
||||
- **vision bins 38–41 (relative progress 0.760–0.840), estimated clip interval 16.22–17.93s** — opposes_predicted_class; evidence: unaligned vision feature rows 242–267; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 33–33 (relative progress 0.660–0.680), estimated clip interval 14.09–14.52s** — opposes_predicted_class; evidence: unaligned vision feature rows 210–216; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.4303 / +0.7252 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.3333 / +0.0765. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.6468.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
As Linn’s associate editor Michael Baadke reports in our November 28 issue, attendees “enthusiastically discussed the topics of growing the hobby, the future of stamp shows, and dealers and philatelic partnerships, along with ways the leading organizations involved in the stamp hobby can work together to make it succeed and grow
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 08
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.712
|
||||
- Predicted-class confidence: 0.812
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/08.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.5797 | 0.888 | +1.1787 | 0.758 |
|
||||
| audio | +0.1601 | 0.090 | +0.1247 | 0.080 |
|
||||
| vision | +0.0394 | 0.022 | -0.2514 | 0.162 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.3202 | -0.2654 |
|
||||
| Text + Vision | +0.0109 | +0.2232 |
|
||||
| Audio + Vision | -0.1553 | -0.0298 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 10–13 (relative progress 0.200–0.280), estimated clip interval 1.02–1.42s** — supports_predicted_class; evidence: brings us
|
||||
- **text bins 17–17 (relative progress 0.340–0.360), estimated clip interval 1.73–1.83s** — supports_predicted_class; evidence: to
|
||||
- **audio bins 46–48 (relative progress 0.920–0.980), estimated clip interval 4.68–4.98s** — opposes_predicted_class; evidence: unaligned audio feature rows 90–96; review the linked source clip at the estimated relative span
|
||||
- **audio bins 42–43 (relative progress 0.840–0.880), estimated clip interval 4.27–4.48s** — opposes_predicted_class; evidence: unaligned audio feature rows 82–86; review the linked source clip at the estimated relative span
|
||||
- **vision bins 23–26 (relative progress 0.460–0.540), estimated clip interval 2.34–2.75s** — opposes_predicted_class; evidence: unaligned vision feature rows 34–39; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 43–43 (relative progress 0.860–0.880), estimated clip interval 4.37–4.48s** — opposes_predicted_class; evidence: unaligned vision feature rows 63–65; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.4306 / +0.8107 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.3601 / +0.0323. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.8594.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
That brings us to tonight, the Universal Design Grand Challenge
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 09
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.848
|
||||
- Predicted-class confidence: 0.918
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/09.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.9453 | 0.940 | -1.5164 | 0.981 |
|
||||
| audio | +0.1079 | 0.052 | -0.0104 | 0.007 |
|
||||
| vision | -0.0172 | 0.008 | +0.0193 | 0.012 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 2.22e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0704 | +0.0883 |
|
||||
| Text + Vision | +0.1085 | -0.0059 |
|
||||
| Audio + Vision | +0.0982 | -0.1718 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 27–30 (relative progress 0.540–0.620), estimated clip interval 2.72–3.13s** — supports_predicted_class; evidence: movie at all
|
||||
- **text bins 35–35 (relative progress 0.700–0.720), estimated clip interval 3.53–3.63s** — supports_predicted_class; evidence: , i
|
||||
- **audio bins 47–49 (relative progress 0.940–1.000), estimated clip interval 4.74–5.04s** — supports_predicted_class; evidence: unaligned audio feature rows 93–98; review the linked source clip at the estimated relative span
|
||||
- **audio bins 2–3 (relative progress 0.040–0.080), estimated clip interval 0.20–0.40s** — supports_predicted_class; evidence: unaligned audio feature rows 3–7; review the linked source clip at the estimated relative span
|
||||
- **vision bins 11–13 (relative progress 0.220–0.280), estimated clip interval 1.11–1.41s** — supports_predicted_class; evidence: unaligned vision feature rows 16–20; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 20–21 (relative progress 0.400–0.440), estimated clip interval 2.02–2.22s** — supports_predicted_class; evidence: unaligned vision feature rows 30–32; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.4802 / +1.6975 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.3075 / +0.1968. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.4469.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
(uhh) I did not like this movie at all, I would not recommend it
|
||||
@@ -0,0 +1,56 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 10
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.260
|
||||
- Predicted-class confidence: 0.862
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/10.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.7953 | 0.665 | -1.3590 | 0.640 |
|
||||
| audio | -0.5743 | 0.213 | +0.6022 | 0.284 |
|
||||
| vision | +0.3315 | 0.123 | -0.1628 | 0.077 |
|
||||
|
||||
Shapley completeness residuals: class -2.22e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.3703 | -0.1735 |
|
||||
| Text + Vision | -0.0413 | -0.0734 |
|
||||
| Audio + Vision | +0.3672 | -0.3802 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 7–9 (relative progress 0.140–0.200), estimated clip interval 1.39–1.99s** — supports_predicted_class; evidence: and you know
|
||||
- **text bins 19–19 (relative progress 0.380–0.400), estimated clip interval 3.78–3.98s** — supports_predicted_class; evidence: see
|
||||
- **audio bins 4–6 (relative progress 0.080–0.140), estimated clip interval 0.80–1.39s** — supports_predicted_class; evidence: unaligned audio feature rows 15–27; review the linked source clip at the estimated relative span
|
||||
- **audio bins 16–16 (relative progress 0.320–0.340), estimated clip interval 3.18–3.38s** — opposes_predicted_class; evidence: unaligned audio feature rows 62–66; review the linked source clip at the estimated relative span
|
||||
- **vision bins 22–26 (relative progress 0.440–0.540), estimated clip interval 4.37–5.37s** — supports_predicted_class; evidence: unaligned vision feature rows 64–79; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.5884 / +2.0317 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.1454 / +0.0604. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.9366.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
(umm) And you know I really do like to see fluffy chick flicks sometimes so I'm not against that but this one was pretty terrible
|
||||
@@ -0,0 +1,57 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 11
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.492
|
||||
- Predicted-class confidence: 0.577
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/11.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +0.7143 | 0.543 | -0.2606 | 0.310 |
|
||||
| audio | +0.1745 | 0.133 | -0.2353 | 0.280 |
|
||||
| vision | -0.4272 | 0.325 | +0.3442 | 0.410 |
|
||||
|
||||
Shapley completeness residuals: class -1.11e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.0550 | +0.1165 |
|
||||
| Text + Vision | -0.0932 | +0.0589 |
|
||||
| Audio + Vision | +0.3224 | -0.2273 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 8–12 (relative progress 0.160–0.260), estimated clip interval 0.93–1.51s** — supports_predicted_class; evidence: would be ashamed
|
||||
- **audio bins 2–4 (relative progress 0.040–0.100), estimated clip interval 0.23–0.58s** — supports_predicted_class; evidence: unaligned audio feature rows 4–11; review the linked source clip at the estimated relative span
|
||||
- **audio bins 33–33 (relative progress 0.660–0.680), estimated clip interval 3.83–3.95s** — supports_predicted_class; evidence: unaligned audio feature rows 75–77; review the linked source clip at the estimated relative span
|
||||
- **vision bins 47–48 (relative progress 0.940–0.980), estimated clip interval 5.46–5.69s** — opposes_predicted_class; evidence: unaligned vision feature rows 80–84; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 4–4 (relative progress 0.080–0.100), estimated clip interval 0.46–0.58s** — opposes_predicted_class; evidence: unaligned vision feature rows 6–8; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.6028 / +0.5872 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -1.1319 / -0.2942. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.5614.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
I would be ashamed to have made this film if I was a director
|
||||
@@ -0,0 +1,56 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 12
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.088
|
||||
- Predicted-class confidence: 0.814
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/12.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.1328 | 0.518 | -0.5134 | 0.367 |
|
||||
| audio | -0.4324 | 0.198 | +0.3258 | 0.233 |
|
||||
| vision | +0.6203 | 0.284 | -0.5603 | 0.400 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.0606 | +0.0851 |
|
||||
| Text + Vision | -0.0521 | +0.0239 |
|
||||
| Audio + Vision | +0.1899 | -0.0598 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 43–45 (relative progress 0.860–0.920), estimated clip interval 11.19–11.97s** — supports_predicted_class; evidence: credit scores is very
|
||||
- **text bins 9–9 (relative progress 0.180–0.200), estimated clip interval 2.34–2.60s** — supports_predicted_class; evidence: had good
|
||||
- **audio bins 2–5 (relative progress 0.040–0.120), estimated clip interval 0.52–1.56s** — opposes_predicted_class; evidence: unaligned audio feature rows 10–30; review the linked source clip at the estimated relative span
|
||||
- **audio bins 10–10 (relative progress 0.200–0.220), estimated clip interval 2.60–2.86s** — opposes_predicted_class; evidence: unaligned audio feature rows 51–56; review the linked source clip at the estimated relative span
|
||||
- **vision bins 10–14 (relative progress 0.200–0.300), estimated clip interval 2.60–3.90s** — supports_predicted_class; evidence: unaligned vision feature rows 38–57; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.3935 / +1.1091 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.0087 / -0.1549. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.1331.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Or worse, an individual previously had good credit, but usually by no fault of their own, or perhaps by fault of their own, they have let their credit sag, and credit scores is very low.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 13
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.421
|
||||
- Predicted-class confidence: 0.525
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/13.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +0.5548 | 0.515 | +0.6172 | 0.544 |
|
||||
| audio | -0.1353 | 0.126 | -0.1868 | 0.165 |
|
||||
| vision | +0.3869 | 0.359 | +0.3305 | 0.291 |
|
||||
|
||||
Shapley completeness residuals: class -1.11e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.1396 | +0.1153 |
|
||||
| Text + Vision | -0.1662 | -0.1338 |
|
||||
| Audio + Vision | -0.2975 | -0.3222 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 19–21 (relative progress 0.380–0.440), estimated clip interval 3.30–3.82s** — supports_predicted_class; evidence: from that data
|
||||
- **text bins 9–10 (relative progress 0.180–0.220), estimated clip interval 1.56–1.91s** — opposes_predicted_class; evidence: could take a
|
||||
- **audio bins 7–9 (relative progress 0.140–0.200), estimated clip interval 1.22–1.74s** — opposes_predicted_class; evidence: unaligned audio feature rows 24–34; review the linked source clip at the estimated relative span
|
||||
- **audio bins 26–27 (relative progress 0.520–0.560), estimated clip interval 4.51–4.86s** — opposes_predicted_class; evidence: unaligned audio feature rows 89–96; review the linked source clip at the estimated relative span
|
||||
- **vision bins 19–21 (relative progress 0.380–0.440), estimated clip interval 3.30–3.82s** — supports_predicted_class; evidence: unaligned vision feature rows 6–7; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 5–5 (relative progress 0.100–0.120), estimated clip interval 0.87–1.04s** — supports_predicted_class; evidence: unaligned vision feature rows 2–2; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: -0.1583 / -0.5077 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.4472 / +0.3904. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: -0.3368.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
For example, I could take a set of data and from that data, I can find a relationship between any two of the given factors or more.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 14
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.526
|
||||
- Predicted-class confidence: 0.582
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/14.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | -0.1328 | 0.115 | +0.2202 | 0.254 |
|
||||
| audio | +0.2039 | 0.177 | +0.1511 | 0.174 |
|
||||
| vision | +0.8151 | 0.708 | +0.4951 | 0.571 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.2578 | -0.0941 |
|
||||
| Text + Vision | -0.2702 | -0.2846 |
|
||||
| Audio + Vision | -0.3436 | -0.2660 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 29–30 (relative progress 0.580–0.620), estimated clip interval 6.00–6.42s** — opposes_predicted_class; evidence: uniform final
|
||||
- **text bins 6–7 (relative progress 0.120–0.160), estimated clip interval 1.24–1.66s** — opposes_predicted_class; evidence: co -
|
||||
- **audio bins 32–35 (relative progress 0.640–0.720), estimated clip interval 6.62–7.45s** — opposes_predicted_class; evidence: unaligned audio feature rows 130–146; review the linked source clip at the estimated relative span
|
||||
- **audio bins 15–15 (relative progress 0.300–0.320), estimated clip interval 3.11–3.31s** — opposes_predicted_class; evidence: unaligned audio feature rows 61–65; review the linked source clip at the estimated relative span
|
||||
- **vision bins 0–2 (relative progress 0.000–0.060), estimated clip interval 0.00–0.62s** — supports_predicted_class; evidence: unaligned vision feature rows 0–9; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 47–48 (relative progress 0.940–0.980), estimated clip interval 9.73–10.14s** — supports_predicted_class; evidence: unaligned vision feature rows 143–149; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: -0.3806 / -0.4348 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.8105 / +0.3206. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: -0.3162.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
He is the co-founder of Rossen and Vettese Limited and the former Executive Director of Uniform Final Examination (UFE) courses at Toronto's York University.
|
||||
@@ -0,0 +1,56 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 15
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +1.053
|
||||
- Predicted-class confidence: 0.893
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/15.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.0820 | 0.496 | +0.7277 | 0.522 |
|
||||
| audio | +0.5320 | 0.244 | +0.2927 | 0.210 |
|
||||
| vision | +0.5664 | 0.260 | +0.3724 | 0.267 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.3240 | -0.1826 |
|
||||
| Text + Vision | -0.2097 | -0.1495 |
|
||||
| Audio + Vision | -0.3615 | -0.3010 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 40–42 (relative progress 0.800–0.860), estimated clip interval 5.87–6.31s** — supports_predicted_class; evidence: power to
|
||||
- **text bins 11–11 (relative progress 0.220–0.240), estimated clip interval 1.61–1.76s** — opposes_predicted_class; evidence: poverty
|
||||
- **audio bins 47–49 (relative progress 0.940–1.000), estimated clip interval 6.89–7.33s** — supports_predicted_class; evidence: unaligned audio feature rows 136–144; review the linked source clip at the estimated relative span
|
||||
- **audio bins 13–14 (relative progress 0.260–0.300), estimated clip interval 1.91–2.20s** — supports_predicted_class; evidence: unaligned audio feature rows 37–43; review the linked source clip at the estimated relative span
|
||||
- **vision bins 11–15 (relative progress 0.220–0.320), estimated clip interval 1.61–2.35s** — supports_predicted_class; evidence: unaligned vision feature rows 24–35; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.1774 / +0.6323 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.7456 / +0.2484. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.6086.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
However, despite their poverty, the family prioritize education because they believed in its power to transform lives
|
||||
@@ -0,0 +1,57 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 16
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -1.836
|
||||
- Predicted-class confidence: 0.914
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/16.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.7042 | 0.876 | -1.2273 | 0.821 |
|
||||
| audio | +0.0968 | 0.050 | -0.0746 | 0.050 |
|
||||
| vision | +0.1450 | 0.075 | -0.1934 | 0.129 |
|
||||
|
||||
Shapley completeness residuals: class -4.44e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0912 | +0.0715 |
|
||||
| Text + Vision | +0.0089 | +0.0956 |
|
||||
| Audio + Vision | -0.0782 | +0.0493 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 43–44 (relative progress 0.860–0.900), estimated clip interval 2.68–2.80s** — supports_predicted_class; evidence: movie
|
||||
- **text bins 15–16 (relative progress 0.300–0.340), estimated clip interval 0.93–1.06s** — supports_predicted_class; evidence: s a
|
||||
- **audio bins 45–49 (relative progress 0.900–1.000), estimated clip interval 2.80–3.11s** — opposes_predicted_class; evidence: unaligned audio feature rows 54–59; review the linked source clip at the estimated relative span
|
||||
- **vision bins 45–46 (relative progress 0.900–0.940), estimated clip interval 2.80–2.93s** — supports_predicted_class; evidence: unaligned vision feature rows 39–41; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 31–32 (relative progress 0.620–0.660), estimated clip interval 1.93–2.06s** — supports_predicted_class; evidence: unaligned vision feature rows 27–29; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.4228 / +1.5463 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.3910 / +0.0861. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.2233.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
It's a terrible, this is a terrible movie
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 17
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.928
|
||||
- Predicted-class confidence: 0.884
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/17.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.8485 | 0.721 | +1.3263 | 0.732 |
|
||||
| audio | -0.2026 | 0.079 | -0.2719 | 0.150 |
|
||||
| vision | +0.5129 | 0.200 | +0.2134 | 0.118 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity -2.22e-16.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0578 | +0.1400 |
|
||||
| Text + Vision | -0.2487 | -0.1171 |
|
||||
| Audio + Vision | -0.1361 | -0.1742 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 29–31 (relative progress 0.580–0.640), estimated clip interval 5.61–6.19s** — supports_predicted_class; evidence: make people
|
||||
- **text bins 21–22 (relative progress 0.420–0.460), estimated clip interval 4.06–4.45s** — supports_predicted_class; evidence: simple,
|
||||
- **audio bins 46–48 (relative progress 0.920–0.980), estimated clip interval 8.89–9.47s** — opposes_predicted_class; evidence: unaligned audio feature rows 175–187; review the linked source clip at the estimated relative span
|
||||
- **audio bins 11–12 (relative progress 0.220–0.260), estimated clip interval 2.13–2.51s** — opposes_predicted_class; evidence: unaligned audio feature rows 42–49; review the linked source clip at the estimated relative span
|
||||
- **vision bins 30–32 (relative progress 0.600–0.660), estimated clip interval 5.80–6.38s** — supports_predicted_class; evidence: unaligned vision feature rows 85–94; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 4–5 (relative progress 0.080–0.120), estimated clip interval 0.77–1.16s** — supports_predicted_class; evidence: unaligned vision feature rows 11–17; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.5239 / +1.5241 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: -0.0036 / -0.1295. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.1456.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
Applying these four design concepts to your presentations is simple, easy and will make people think you turned into a design guru.
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 18
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **negative**
|
||||
- Intensity: -0.171
|
||||
- Predicted-class confidence: 0.490
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/18.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | -0.3800 | 0.380 | +0.5236 | 0.558 |
|
||||
| audio | +0.1779 | 0.178 | +0.0299 | 0.032 |
|
||||
| vision | +0.4407 | 0.441 | -0.3844 | 0.410 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0407 | +0.1208 |
|
||||
| Text + Vision | +0.2032 | +0.0182 |
|
||||
| Audio + Vision | +0.0530 | -0.1900 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 3–6 (relative progress 0.060–0.140), estimated clip interval 1.13–2.63s** — opposes_predicted_class; evidence: in denmark - the
|
||||
- **text bins 15–15 (relative progress 0.300–0.320), estimated clip interval 5.64–6.02s** — opposes_predicted_class; evidence: – certified
|
||||
- **audio bins 7–10 (relative progress 0.140–0.220), estimated clip interval 2.63–4.14s** — supports_predicted_class; evidence: unaligned audio feature rows 52–81; review the linked source clip at the estimated relative span
|
||||
- **audio bins 23–23 (relative progress 0.460–0.480), estimated clip interval 8.65–9.03s** — opposes_predicted_class; evidence: unaligned audio feature rows 171–178; review the linked source clip at the estimated relative span
|
||||
- **vision bins 43–46 (relative progress 0.860–0.940), estimated clip interval 16.17–17.68s** — supports_predicted_class; evidence: unaligned vision feature rows 240–263; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 10–10 (relative progress 0.200–0.220), estimated clip interval 3.76–4.14s** — supports_predicted_class; evidence: unaligned vision feature rows 56–61; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: -0.4081 / +0.0710 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +1.4230 / +0.5937. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.0147.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
-And in Denmark - the first Baltic Cod fishery has been MSC – certified -Meanwhile, the Faeroese Mackerel Fishery has been denied MSC certification based on the fact that the fishery has failed to reach an agreement on mackerel quotas with Norway and the European Union.
|
||||
@@ -0,0 +1,56 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 19
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.047
|
||||
- Predicted-class confidence: 0.474
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/19.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | -0.0595 | 0.080 | +0.1895 | 0.273 |
|
||||
| audio | -0.0395 | 0.053 | -0.1530 | 0.221 |
|
||||
| vision | +0.6449 | 0.867 | +0.3506 | 0.506 |
|
||||
|
||||
Shapley completeness residuals: class -1.11e-16, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | -0.0337 | +0.2002 |
|
||||
| Text + Vision | +0.0045 | -0.0336 |
|
||||
| Audio + Vision | -0.1129 | -0.1968 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 34–36 (relative progress 0.680–0.740), estimated clip interval 9.30–10.13s** — opposes_predicted_class; evidence: all over imperfections
|
||||
- **text bins 12–13 (relative progress 0.240–0.280), estimated clip interval 3.28–3.83s** — opposes_predicted_class; evidence: own up to
|
||||
- **audio bins 16–18 (relative progress 0.320–0.380), estimated clip interval 4.38–5.20s** — opposes_predicted_class; evidence: unaligned audio feature rows 87–103; review the linked source clip at the estimated relative span
|
||||
- **audio bins 0–1 (relative progress 0.000–0.040), estimated clip interval 0.00–0.55s** — opposes_predicted_class; evidence: unaligned audio feature rows 0–10; review the linked source clip at the estimated relative span
|
||||
- **vision bins 26–30 (relative progress 0.520–0.620), estimated clip interval 7.12–8.48s** — supports_predicted_class; evidence: unaligned vision feature rows 106–126; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.0467 / +0.1815 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.2694 / +0.1343. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +0.1936.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
People are surprisingly forgiving brands when they own up to mistakes, and unfortunately some haters out there love to point fingers and jump all over imperfections, but for the most part, people understand
|
||||
@@ -0,0 +1,58 @@
|
||||
# Q3 explanation card — E2_MoFE_Shapley — 20
|
||||
|
||||
## Prediction
|
||||
|
||||
- Polarity: **positive**
|
||||
- Intensity: +0.760
|
||||
- Predicted-class confidence: 0.835
|
||||
- Source clip: 附件4-可解释专项视频样本与特征文件/附件4-可解释专项视频样本与特征文件/未对齐版本/videos/20.mp4
|
||||
|
||||
## Exact modality Shapley
|
||||
|
||||
Positive values support the predicted class logit; negative values oppose it. Shares use absolute values and are model decision contributions, not real-world emotion importance.
|
||||
|
||||
| Modality | Class logit contribution | Absolute share | Intensity contribution | Absolute share |
|
||||
|---|---:|---:|---:|---:|
|
||||
| text | +1.8305 | 0.936 | +1.4182 | 0.756 |
|
||||
| audio | -0.0497 | 0.025 | -0.3878 | 0.207 |
|
||||
| vision | +0.0764 | 0.039 | +0.0693 | 0.037 |
|
||||
|
||||
Shapley completeness residuals: class 0.00e+00, intensity 0.00e+00.
|
||||
## Pairwise Shapley interaction
|
||||
|
||||
| Pair | Class logit | Intensity |
|
||||
|---|---:|---:|
|
||||
| Text + Audio | +0.1485 | +0.4687 |
|
||||
| Text + Vision | -0.1556 | -0.1111 |
|
||||
| Audio + Vision | -0.1757 | -0.2182 |
|
||||
|
||||
## Local evidence segments
|
||||
|
||||
Local counterfactual scores are the predicted-class logit difference after hiding a 1/3/5-bin window, averaged equally across the three scales. E1 ranks its router utility and is evaluated separately.
|
||||
|
||||
- **text bins 11–13 (relative progress 0.220–0.280), estimated clip interval 2.41–3.07s** — supports_predicted_class; evidence: of the description of
|
||||
- **text bins 28–28 (relative progress 0.560–0.580), estimated clip interval 6.13–6.35s** — supports_predicted_class; evidence: updates and
|
||||
- **audio bins 33–34 (relative progress 0.660–0.700), estimated clip interval 7.23–7.67s** — opposes_predicted_class; evidence: unaligned audio feature rows 142–151; review the linked source clip at the estimated relative span
|
||||
- **audio bins 26–27 (relative progress 0.520–0.560), estimated clip interval 5.70–6.13s** — supports_predicted_class; evidence: unaligned audio feature rows 112–120; review the linked source clip at the estimated relative span
|
||||
- **vision bins 13–15 (relative progress 0.260–0.320), estimated clip interval 2.85–3.50s** — opposes_predicted_class; evidence: unaligned vision feature rows 42–51; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
- **vision bins 34–35 (relative progress 0.680–0.720), estimated clip interval 7.45–7.89s** — opposes_predicted_class; evidence: unaligned vision feature rows 110–116; candidate frame time is estimated from relative progress
|
||||
- Candidate frame: 
|
||||
|
||||

|
||||
|
||||
## Faithfulness checks
|
||||
|
||||
- Comprehensiveness after deleting the top 10% / 30% cells: +0.6271 / +1.7907 predicted-class logit.
|
||||
- Sufficiency gap when retaining the top 10% / 30%: +0.1621 / -0.0744. Smaller absolute gaps are better.
|
||||
- Mean deletion logit drop over 0–70% deletion: +1.4280.
|
||||
|
||||
## Provenance limit
|
||||
|
||||
Attachment 4 supplies unaligned feature sequences without word/audio/frame timestamps. Feature rows are traced to source rows and normalized progress. Clip-time estimates multiply that progress by the video duration; they are approximate review locations, not physical alignment timestamps.
|
||||
|
||||
Occlusion and Shapley values describe this trained model's response to masked inputs. They do not establish causal effects or prove the emotion expressed by a person.
|
||||
|
||||
## Transcript
|
||||
|
||||
And of course, click in the link of the description of this video for more, and we'll have more live updates and a stock market video (wrap-up) at the end of the day today.
|
||||
Reference in New Issue
Block a user