Aaquib Syed

College Park, MD

Bio

Updated 03/22/26

Aaquib Syed is a CS and Mathematics undergraduate student at the University of Maryland, College Park, where he is a Banneker-Key Scholar. He was a fellow in the MATS 5.0 program under Neel Nanda's supervision, conducting mechanistic interpretability research on how refusal is implemented in large language models. His most notable work, "Refusal in Language Models Is Mediated by a Single Direction" (NeurIPS 2024), co-authored with Andy Arditi and others, showed that refusal behavior across 13 open-source chat models is controlled by a single direction in the residual stream. He also co-authored "Attribution Patching Outperforms Automated Circuit Discovery" (NeurIPS 2024 ATTRIB workshop) and work on mechanistic unlearning and model pruning. He is currently a Student Researcher at Google DeepMind on the Frontier Safety team, where he focuses on evaluating and forecasting dangerous AI capabilities.

Community Signal

Updated 03/22/26

0Upvotes

0Downvotes

0Endorsements

No endorsements yet.

Grants

Updated 03/22/26

LTFF 2024 Q1 - Aaquib Syed

from Long-Term Future Fundfunds.effectivealtruism.org

recipient$30,000

Aaquib Syed

Bio

Community Signal

Links

Grants