Weekly Schedule
The schedule along with the list of topics to be covered are tentative and subject to change.
| Week | Date | Module | Topic | Reading | Assignment | |
|---|---|---|---|---|---|---|
| 1 | T | 8/25 | Introduction | How AI fails, and what “trust” decomposes into (slides) | Post a note on Teams to introduce yourself. Review the Syllabus. | |
| R | 8/27 | Introduction | Convolutional networks and transformers (notebook) | Phuong & Hutter, Formal Algorithms for Transformers | Setup software and development environment for the course (see here) Assignment 0 out | |
| 2 | T | 9/1 | No class: Ravi traveling | |||
| R | 9/3 | No class: Ravi traveling | ||||
| 3 | T | 9/8 | Introduction | Transformer review and post-training (notebook) | Rafailov et al., Direct Preference Optimization (§3-4) | |
| R | 9/10 | Robustness Attacks | Adversarial examples as constrained optimization (notebook) | Goodfellow et al., Explaining and Harnessing Adversarial Examples; Madry et al., Towards Deep Learning Models Resistant to Adversarial Attacks (§2); Gilmer et al., Motivating the Rules of the Game for Adversarial Example Research | ||
| 4 | T | 9/15 | Robustness Attacks | Jailbreaking: GCG as discrete PGD (notebook) | Zou et al., Universal and Transferable Adversarial Attacks on Aligned Language Models | Assignment 0 due Assignment 1 out |
| R | 9/17 | Robustness Defenses | Empirical defenses, on images and on language models (notebook) | Madry et al., Towards Deep Learning Models Resistant to Adversarial Attacks; Sharma et al., Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming | ||
| 5 | T | 9/22 | Robustness Defenses | Empirical defenses, continued: evaluating a defense (notebook) | Athalye et al., Obfuscated Gradients Give a False Sense of Security; Carlini et al., On Evaluating Adversarial Robustness (§2-3) | |
| R | 9/24 | Robustness Guarantees | Certified robustness: bound propagation (notebook) | Gowal et al., On the Effectiveness of Interval Bound Propagation for Training Verifiably Robust Models; Singh et al., An Abstract Domain for Certifying Neural Networks (skim) | ||
| 6 | T | 9/29 | Robustness Guarantees Interpretability | Certified robustness, continued (notebook) Input attribution (notebook) | Sundararajan et al., Axiomatic Attribution for Deep Networks | Assignment 1 due Assignment 2 out |
| R | 10/1 | Interpretability | Input attribution and its weaknesses, continued (notebook) | Adebayo et al., Sanity Checks for Saliency Maps | ||
| 7 | T | 10/6 | No class: Ravi traveling | |||
| R | 10/8 | Interpretability | Probes, concepts, and steering | |||
| 8 | T | 10/13 | Interpretability | Circuits and sparse autoencoders | ||
| R | 10/15 | Review and catch-up | Assignment 2 due | |||
| 9 | T | 10/20 | Midterm | |||
| R | 10/22 | Training-time Failures | Data poisoning and backdoors | Assignment 3 out | ||
| 10 | T | 10/27 | Training-time Failures | Sleeper agents and emergent misalignment | ||
| R | 10/29 | Training-time Failures | Reward hacking and specification gaming | |||
| 11 | T | 11/3 | Calibration and Evaluation | Does the model know when it is wrong? | Project proposal due | |
| R | 11/5 | Calibration and Evaluation | Benchmarks, red teams, and judges | |||
| 12 | T | 11/10 | Guest lecture: Provenance and watermarking | |||
| R | 11/12 | Agents | Agent architecture and the agent threat model | |||
| 13 | T | 11/17 | Agents | Prompt injection | Assignment 3 due Assignment 4 out | |
| R | 11/19 | Agents | Structural defenses | |||
| 14 | T | 11/24 | No class: Fall recess | |||
| R | 11/26 | No class: Fall recess | ||||
| 15 | T | 12/1 | Agents | Runtime oversight | Assignment 4 due 12/5 | |
| R | 12/3 | No class: Ravi traveling | ||||
| 16 | T | 12/8 | Wrap-up | Project presentations | ||
| R | 12/10 | Wrap-up | What we did not cover | Project report due | ||
| Finals | W | 12/16 | Final Exam in CSB 130 (6:20-8:20 pm) |