Weekly Schedule

The schedule along with the list of topics to be covered are tentative and subject to change.

Week DateModuleTopicReadingAssignment
1T8/25IntroductionHow AI fails, and what “trust” decomposes into (slides) Post a note on Teams to introduce yourself.
Review the Syllabus.
 R8/27IntroductionConvolutional networks and transformers (notebook)Phuong & Hutter, Formal Algorithms for TransformersSetup software and development environment for the course (see here)
Assignment 0 out
2T9/1 No class: Ravi traveling  
 R9/3 No class: Ravi traveling  
3T9/8IntroductionTransformer review and post-training (notebook)Rafailov et al., Direct Preference Optimization (§3-4) 
 R9/10Robustness AttacksAdversarial examples as constrained optimization (notebook)Goodfellow et al., Explaining and Harnessing Adversarial Examples; Madry et al., Towards Deep Learning Models Resistant to Adversarial Attacks (§2); Gilmer et al., Motivating the Rules of the Game for Adversarial Example Research 
4T9/15Robustness AttacksJailbreaking: GCG as discrete PGD (notebook)Zou et al., Universal and Transferable Adversarial Attacks on Aligned Language ModelsAssignment 0 due
Assignment 1 out
 R9/17Robustness DefensesEmpirical defenses, on images and on language models (notebook)Madry et al., Towards Deep Learning Models Resistant to Adversarial Attacks; Sharma et al., Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming 
5T9/22Robustness DefensesEmpirical defenses, continued: evaluating a defense (notebook)Athalye et al., Obfuscated Gradients Give a False Sense of Security; Carlini et al., On Evaluating Adversarial Robustness (§2-3) 
 R9/24Robustness GuaranteesCertified robustness: bound propagation (notebook)Gowal et al., On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models; Singh et al., An Abstract Domain for Certifying Neural Networks (skim) 
6T9/29Robustness Guarantees
Interpretability
Certified robustness, continued (notebook)
Input attribution (notebook)
Sundararajan et al., Axiomatic Attribution for Deep NetworksAssignment 1 due
Assignment 2 out
 R10/1InterpretabilityInput attribution and its weaknesses, continued (notebook)Adebayo et al., Sanity Checks for Saliency Maps 
7T10/6 No class: Ravi traveling  
 R10/8InterpretabilityProbes, concepts, and steering  
8T10/13InterpretabilityCircuits and sparse autoencoders  
 R10/15 Review and catch-up Assignment 2 due
9T10/20 Midterm  
 R10/22Training-time FailuresData poisoning and backdoors Assignment 3 out
10T10/27Training-time FailuresSleeper agents and emergent misalignment  
 R10/29Training-time FailuresReward hacking and specification gaming  
11T11/3Calibration and EvaluationDoes the model know when it is wrong? Project proposal due
 R11/5Calibration and EvaluationBenchmarks, red teams, and judges  
12T11/10 Guest lecture: Provenance and watermarking  
 R11/12AgentsAgent architecture and the agent threat model  
13T11/17AgentsPrompt injection Assignment 3 due
Assignment 4 out
 R11/19AgentsStructural defenses  
14T11/24 No class: Fall recess  
 R11/26 No class: Fall recess  
15T12/1AgentsRuntime oversight Assignment 4 due 12/5
 R12/3 No class: Ravi traveling  
16T12/8Wrap-upProject presentations  
 R12/10Wrap-upWhat we did not cover Project report due
FinalsW12/16 Final Exam in CSB 130 (6:20-8:20 pm)