DS 4440  ·  Northeastern University  ·  Khoury College of Computer Sciences
Neural Networks  ·  Fall 2026

Course Description

This course is a hands-on introduction to modern neural network models. We will cover the fundamentals of neural networks, and work our way up to modern generative architectures. We will cover stochastic gradient descent and backpropagation, along with related fitting techniques.

A significant component of this course will focus on interpretability: understanding how generative models manage to realize the functionality they offer. We will make use of the National Deep Inference Fabric (NDIF), a Northeastern-led project.

Logistics

I will assign homeworks to be done without AI, but because I cannot police this, they will not be graded. Instead, they are meant to help you learn the material and prep for the brief in-class assessments we will have regularly; these are intended to be straightforward if you have done the homeworks.

Grading

35%In-class quizzes and exercises
25%Midterm
40%Final project

Your lowest quiz score will be dropped.

Prerequisites

Prior exposure to machine learning is recommended. Working knowledge of Python is required. Familiarity with linear algebra, basic calculus, and probability will be largely assumed, though we will review key prerequisites.

Midterm

An in-class midterm covering the foundations of deep learning.

Projects

A major component of the course is a project completed in pairs. Projects should concern some aspect of model interpretability and will use NDIF.

See project guidelines here

Academic Integrity and AI

No AI will be allowed while taking the in-class quizzes or midterm (all closed book, as well). You can use AI freely on homeworks (though these won't be graded and we recommend minimal use to ensure you understand the material) and on projects (but you must understand the code).

See the Northeastern Academic Integrity Policy for general academic integrity policy guidelines (which apply here).

Schedule

Dates and materials will be updated as the semester progresses. Readings reference d2l unless otherwise noted.

Date Topic Readings Notes Materials
9/9 (W) Course aims, logistics; Review of supervised learning / Perceptron d2l: Introduction; Original Perceptron (1957) Join Piazza Slides; Perceptron notebook; Lecture notes: scribbled; Lecture notes: typeset (by Claude)
9/14 (M) Logistic Regression and Optimization via SGD d2l: Preliminaries Linear models notebook; Lecture notes: scribbled; Lecture notes: typeset (by Claude)
9/16 (W) Beyond Linear Models: The Multi-Layer Perceptron d2l: MLPs (4.1) Quiz 1; based on HW 1 MLP notebook; Lecture notes: scribbled; Lecture notes: typeset (by Claude); Quiz 1 solutions
9/21 (M) Abstractions: Layers and Computation Graphs d2l: Layers and blocks Computation graphs notebook; Lecture notes: scribbled; Lecture notes: typeset (by Claude)
9/23 (W) Backpropagation I d2l: Autodiff; Colah's blog; Rumelhart et al. (1986) Backprop notebook; Lecture notes: scribbled; Lecture notes: typeset (by Claude)
9/28 (M) Backpropagation II d2l: Backprop Custom layer notebook; Lecture notes: scribbled; Lecture notes: typeset (by Claude)
9/30 (W) Optimizer matters: Training Neural Networks in Practice d2l: Optimization Quiz 2; based on HW 2 Optimization notebook; Lecture notes: scribbled; Lecture notes: typeset (by Claude)
10/5 (M) Learning representations of discrete things: Embeddings d2l: Word embeddings (14.1)
10/7 (W) Convolutional Neural Networks (CNNs) d2l: CNNs (6.1)
10/12 (M) No class — Indigenous Peoples Day
10/14 (W) Stacking ConvNets, residual connections, and other tricks d2l: CNNs (6.2–6.5); Modern CNNs (7.1, 7.5–7.7) Quiz 3; based on HW 3
10/19 (M) Recurrent Neural Networks I d2l: RNNs (8.1, 8.4); Karpathy: Unreasonable Effectiveness of RNNs
10/21 (W) Recurrent Neural Networks II d2l: RNNs (8.7)
10/26 (M) Transformers, self-supervision, contextualized embeddings d2l, Ch. 11
10/28 (W) More Transformers; BERT and BERTology Quiz 4; based on HW 4
11/2 (M) Midterm review
11/4 (W) Midterm
11/9 (M) Posttraining (instruction tuning and RLHF) 1
11/11 (W) No class — Veterans Day
11/16 (M) Posttraining (instruction tuning and RLHF) 2 Project proposals due
11/18 (W) How do LLMs work? Interpretability
11/23 (M) Guest lecture: Activation verbalization
11/25 (W) No class — Fall break
11/30 (M) Diffusion Models Step-by-Step Diffusion: An Elementary Tutorial Quiz 5; based on HW 5
12/2 (W) Ethical problems with generative models
12/7 (M) Dedicated project feedback and help
12/9 (W) Project presentations