Noam Razin

Noam Razin

Assistant Professor

Computer Science & AI DepartmentBar-Ilan University

razinno [at] cs.biu.ac.il

I am an Assistant Professor in the Computer Science & AI Department at Bar-Ilan University. My research focuses on the fundamentals of deep learning and modern artificial intelligence (AI) systems. By combining mathematical analysis with systematic experimentation, I aim to develop theories that shed light on how modern AI works, identify potential failures, and yield principled methods for improving efficiency, reliability, and performance.

Previously, I was a postdoctoral fellow at Princeton University, hosted by Sanjeev Arora. Before that, I obtained my PhD in Computer Science at Tel Aviv University under the supervision of Nadav Cohen.

I am recruiting MSc and PhD students. Please see the note below if you are interested in working with me.

Recent Research

Recently, I have been working on aspects of converting language models to useful AI systems (aka post-training), including failures of preference-based alignment [1] and policy gradient methods [2], what makes a good proxy reward function [3, 4], catastrophic forgetting [5], and reward model generalization [6].

News

Publications

See also Google Scholar* indicates equal contribution

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

Shuning Shang*, Hubert Strauss*, Stanley Wei, Sanjeev Arora, Noam Razin

arXiv:2604.25872, 2026

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting

Howard Chen, Noam Razin, Karthik Narasimhan, Danqi Chen

International Conference on Machine Learning (ICML), 2026

Why is Your Language Model a Poor Implicit Reward Model?

Noam Razin, Yong Lin, Jiarui Yao, Sanjeev Arora

International Conference on Learning Representations (ICLR), 2026

What Makes a Reward Model a Good Teacher? An Optimization Perspective

Noam Razin, Zixuan Wang, Hubert Strauss, Stanley Wei, Jason D. Lee, Sanjeev Arora

Advances in Neural Information Processing Systems (NeurIPS), 2025

The Implicit Bias of Structured State Space Models Can Be Poisoned With Clean Labels

Yonatan Slutzky*, Yotam Alexander*, Noam Razin, Nadav Cohen

Advances in Neural Information Processing Systems (NeurIPS), 2025

Selected Talks

  • What Makes a Good Proxy Reward Function for Language Model Post-Training?

    HUJI Machine Learning Club · Jun 2026

  • Why is Your Language Model a Poor Implicit Reward Model?

    Princeton Language and Intelligence Seminar · Oct 2025

  • Understanding and Overcoming Pitfalls in Language Model Alignment

    EPFL AI Fundamentals Seminar · Sep 2025

  • Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

    Deep Learning: Classics and Trends Seminar · Jan 2025

  • Analyses of Policy Gradient for Language Model Finetuning and Optimal Control

    MPI MiS + UCLA Math Machine Learning Seminar · Mar 2024

Teaching

Current

  • Theoretical Foundations of Deep Learning

    Lecturer · Bar-Ilan University · 2026–

Past

  • Fundamentals of Deep Learning

    Guest Lecturer · Princeton University · 2025

  • Introduction to Reinforcement Learning

    Guest Lecturer · Princeton University · 2025

  • First Steps in Research Honors Seminar

    Guest Lecturer · Tel Aviv University · 2021–2024

  • Foundations of Deep Learning

    Teaching Assistant · Tel Aviv University · 2021–2023