Stanford CS 329M Lectures (September - December 2026):
[Week 01] Sept. 22 / Lecture 1 - Fundamentals of Machine Programming
[Week 02] Sept. 29 / Lecture 2 - Tour of the State of MP (e.g., Claude Code, GitHub Copilot, Cursor, Devin Desktop (Windsurf), Gemini Code Assist, Lovable, Replit Agent, etc.)
[Week 03] Oct. 6 / Lecture 3 - Core Tenets & First Principles of ML / MP, Deeper Dive into 4-3-2 of MP
[Week 04] Oct.13 / Lecture 4 - Principles of Data: Ground Truth, Semi-Trust, and Static / Continuous Learning
[Week 05] Oct. 20 / Lecture 5 - Meta-Data, Meta-Reasoning, Latent Spaces & Higher-Dimensional Analyses
[Week 06] Oct. 27 / Lecture 6 - Performance Implications of MP - Cloud, Local, Hybrid, and Other
[Week 07] Nov. 3 / Stanford Democracy Day (Classes Canceled)
[Week 08] Nov. 10 / Lecture 7 - Classical & Future Software Engineering (SE) with Temporal Reasoning
[Week 09] Nov. 17 / Lectures 8 - Future of Programming & Programming Languages (HIPLs, MIPLs, Intentional PLs, etc.)
[Week 10] Nov. 24: Autumn Break (No Classes)
[Week 11] Dec. 1 / Lecture 9 - Future of Machine Programming & the Criticality of the Next Decade
Course Description
The field of machine programming (MP) is concerned with the automation of software development. Given the recent advances in software algorithms, hardware efficiency and capacity, and an ever increasing availability of code data, it is now possible to train machines to help develop software. In this course, we teach students how to build real-world MP systems. We begin with a high-level overview of the field, including an abbreviated analysis of state-of-the-art (e.g., code assistants). Next, we discuss the foundations of MP and the key areas for innovation, some of which are unique to MP. We close with a discussion of current limitations and future directions of MP. This course includes a nine-week hands-on project, where students (as individuals or in a small group) will create their own MP system and demonstrate it to the class.
While some overlap exists between traditional techniques to train machines to perform non-programming tasks (e.g., natural language processing, computer vision, etc.), teaching machines to perform programming-specific tasks has uniqueness in at least two dimensions. First, there are certain techniques that are more (or less) effective for MP, such as using self-supervision to learn from the large corpora of unlabeled open-source code. Second, software reasoning is fundamentally multi-dimensional; that is, there exist multiple unique ways to learn from software (e.g., static analysis, dynamic analysis, input/output specifications, program state reinforced-convergence, hardware telemetric data, etc.). In this course, we discuss each of these techniques (and others) and how they can be effectively applied to MP systems.
This course is primarily intended for Stanford MS and PhD graduate students. However, advanced (senior-level) and committed undergraduates have successfully completed this course. Hard-working and disciplined students have historically done well in CS 329M even if they are relatively new to ML, PL, SE, and systems.
Examples of Notable Prior CS 329M Projects:
The following short list of papers are exemplars of CS 329M student projects. Each target a different subdomain within the field of MP. Even though they are diverse from each other, they represent a small fraction (< 1%) of all possible domains you can target for your CS 329M project.
"Un-reinventing the Wheel: Large Scale Open Source Code Search" (slides, Hellman and Spector)
"Stylish Bot: a machine learning and rule-based approach to provide style feedback on students' coding assignments" (Chuen)
"Evaluating and Improving the Security of AI-Based Code Assistants" (slides, Gu)
"Side Effect Detection through Meta-reasoning and Learning" (Li and Zhu)