ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing
Deep neural networks often exploit spurious associations in their training data, a failure known as shortcut learning. Concept-based explainability methods screen for shortcuts by testing whether concepts such as a patient's sex or scanner settings can be decoded from a network layer. Because each concept is evaluated in isolation, these methods can mistake correlations between concepts as evidence that the model uses them. We introduce ICON decomposition, which instead quantifies how much of a layer's variance each concept explains after accounting for all other concepts and the outcome. On synthetic data with known ground truth, ICON recovers concept importance more accurately than seven alternative baseline methods. On skin-lesion and brain-imaging models, it isolates the concepts on which a model genuinely relies, quantifies the representation unexplained by any of the supplied concepts, and yields sparse explanations that we validate by retraining and out-of-distribution testing.
C1 reading
Select any word for its Thai meaning and pronunciation.
แปลไทยทั้งบท
เครือข่ายประสาทลึกมักจะใช้สมาคมที่ซับซ้อนในข้อมูลการฝึกอบรมของพวกเขา ซึ่งเป็นการล้มเหลวที่รู้จักกันในชื่อการเรียนรู้ทางสั้น. วิธีการอธิบายที่ขึ้นอยู่กับแนวคิดตรวจสอบทางล่วงทางโดยการทดสอบว่าแนวคิด เช่น ความเพศของผู้ป่วยหรือการตั้งค่าสแกนเนอร์ทำได้ออกรหัสได้จากชั้นเครือข่ายหรือไม่. เพราะแนวคิดแต่ละแนวคิดถูกประเมินแยกกัน โดยวิธีเหล่านี้ทำได้ทำให้ความสัมพันธ์ระหว่างแนวคิดผิดพลาดเป็นหลักฐานว่าตัวอย่างนั้นใช้มัน.
เรานําเสนอความละลาย ICON ซึ่งแทนจะระบุปริมาณของความแตกต่างของชั้นแต่ละแนวคิดที่อธิบายหลังจากคํานวณสำหรับแนวคิดอื่น ๆ และผล. จากข้อมูลสังเคราะห์ที่มีความเป็นจริงในพื้นที่ที่รู้จัก ICON ยกคืนความสําคัญของแนวคิดได้อย่างแม่นยํากว่าวิธีเบาซไลน์ทางเลือกเจ็ดวิธี. ในรูปแบบบาดเจ็บผิวหนัง และรูปภาพสมอง มันแยกแยกแนวคิดที่รูปแบบนั้นพึ่งพากันอย่างแท้จริง มันปริมาณการแสดงตัวที่ไม่ได้ถูกอธิบายโดยแนวคิดใด ๆ ที่นํามาให้ และผลิตการอธิบายที่หายาก ที่เรายืนยันโดยการฝึกอบรมใหม่ และการทดสอบนอกการจําหน่าย.
ประโยคและวลีที่ใช้ได้จริงจากเรื่องนี้
Useful phrases from this story
การตรวจสอบวิธีการอธิบาย.
From the storyConcept-based explainability methods screen for shortcuts by testing whether concepts such as a patient's sex or scanner settings can be decoded from a network layer.
การทดสอบว่า แนวคิด เช่น.
From the storyConcept-based explainability methods screen for shortcuts by testing whether concepts such as a patient's sex or scanner settings can be decoded from a network layer.
สามารถออกรหัสได้จาก.
From the storyConcept-based explainability methods screen for shortcuts by testing whether concepts such as a patient's sex or scanner settings can be decoded from a network layer.
การประเมินแยกแยก.
From the storyBecause each concept is evaluated in isolation, these methods can mistake correlations between concepts as evidence that the model uses them.
สามารถผิดพลาดความสัมพันธ์ระหว่างแนวคิด.
From the storyBecause each concept is evaluated in isolation, these methods can mistake correlations between concepts as evidence that the model uses them.
Save & Review
Only words saved from this story appear here.