Shu Yang 杨澍
About me
I am a Ph.D. student in Computer Science at King Abdullah University of Science and Technology (KAUST), advised by Prof. Di Wang in the PRADA Lab (Provable Responsible AI and Data Analytics Lab). I received my M.Sc. in Computational Linguistics from the University of Macau, advised by Prof. Derek F. Wong. I have been a visiting scholar at the University of Edinburgh and Tsinghua University. I am currently a Visiting Student Researcher at the Stanford Digital Economy Lab, Stanford HAI, where I work closely with Sandy Pentland and Jiaxin Pei.
We are building the next generation of human–agent collaboration. Today's AI assistants are private, single-user tools, while real work happens in shared conversations. With Kordi, an open-source, AI-native collaboration workspace, we treat agents as participants in the conversation: people and their agents share the same space, and anyone can call on any agent right where the context lives. This builds on my earlier work on LLM alignment and monitorability, including sycophancy, reward hacking, and reasoning-induced misalignment.
News
- 2026.09 Hearing is Believing? Evaluating and Analyzing Audio Language Model Sycophancy with SYAUDIO accepted by NeurIPS 2026.
- 2026.08 Four papers accepted by EMNLP 2026:
- LLMs Regret Before They Say It: Early Detection and Compositional Architecture of Regret in Hidden States
- Understanding and Mitigating Cross-lingual Privacy Leakage via Language-specific and Universal Privacy Neurons
- Neuron-Guided Fine-Tuning: Unlocking Efficient Alignment Mechanisms for Large Language Models
- STARS: Skill Triage via Activation Risk Scoring for LLM Agents
- 2026.07 Can AI Truly Represent Your Voice in Deliberations? A Comprehensive Study of Large-Scale Opinion Aggregation with LLMs accepted by CoLM 2026.
- 2026.05 Joined the Stanford Digital Economy Lab at Stanford HAI as a Visiting Student Researcher.
- 2026.04 Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models accepted by ICML 2026.
- 2026.04 Five papers accepted by ACL 2026:
- AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
- Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images Oral
- Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
- JARVIS or Ultron? A Survey on the Safety and Security Threats of Computer-Using Agents
- 2026.02 My first blog, Misalignments and RL failure modes in the early stage of superintelligence, accepted by the ICLR 2026 Blogpost Track.
- 2026.01 Four papers accepted by ICLR 2026:
- 2026.01 Understanding Aha Moments: from External Observations to Internal Mechanisms accepted by TACL.
- 2025.12 The Web Tool Trap: Understanding and Mitigating Over-Reliance in LLM Browsing Agents accepted as an extended abstract by AAMAS 2026.
- 2025.11 When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models accepted by AAAI 2026.
- 2025.10 Our workshop proposal for the First Workshop on LLM Persona Modeling was accepted by NeurIPS 2025. See you in Mexico City!
- 2025.09 Began my visit at EdinburghNLP, University of Edinburgh.
- 2025.09 EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification accepted by NeurIPS 2025.
- 2025.08 Two papers accepted by EMNLP 2025:
- 2025.08 RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns accepted by TACL.
- 2025.05 Three papers accepted by ACL 2025:
- 2025.04 Started my visit at the THUNLP group, Tsinghua University.
- 2024.11 Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore accepted by COLING 2025.
- 2024.11 Our survey A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions accepted by Computational Linguistics.
- 2024.10 DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios accepted by the NeurIPS 2024 Datasets & Benchmarks Track.
- 2024.08 Began my Ph.D. journey at the PRADA Lab, KAUST, under the supervision of Prof. Di Wang.
- 2024.07 Model Autophagy Analysis to Explicate Self-consumption within Human-AI Interactions accepted by CoLM 2024.
- 2024.06 Completed my Master's degree at the University of Macau, under the supervision of Prof. Derek Wong.
Publications
Selected conference and journal papers. See Google Scholar for the full list.
2026
- NeurIPS 2026 Hearing is Believing? Evaluating and Analyzing Audio Language Model Sycophancy with SYAUDIO
- EMNLP 2026 LLMs Regret Before They Say It: Early Detection and Compositional Architecture of Regret in Hidden States
- EMNLP 2026 Understanding and Mitigating Cross-lingual Privacy Leakage via Language-specific and Universal Privacy Neurons
- EMNLP 2026 Neuron-Guided Fine-Tuning: Unlocking Efficient Alignment Mechanisms for Large Language Models
- EMNLP 2026 STARS: Skill Triage via Activation Risk Scoring for LLM Agents
- CoLM 2026 Can AI Truly Represent Your Voice in Deliberations? A Comprehensive Study of Large-Scale Opinion Aggregation with LLMs
- ICML 2026 Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
- ACL 2026 AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
- ACL 2026 Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
- ACL 2026 Oral Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images
- ACL 2026 Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
- ACL 2026 JARVIS or Ultron? A Survey on the Safety and Security Threats of Computer-Using Agents
- ICLR 2026 When Thinking Backfires: Mechanistic Insights into Reason-induced Misalignment
- ICLR 2026 Neuron-Aware Data Selection in Instruction Tuning for Large Language Models
- ICLR 2026 Evaluating Data Influence in Meta Learning
- ICLR 2026 Dissecting Representation Misalignment in Contrastive Learning via Influence Function
- ICLR 2026 Blogpost Misalignments and RL failure modes in the early stage of superintelligence
- TACL Understanding Aha Moments: from External Observations to Internal Mechanisms
- AAMAS 2026 The Web Tool Trap: Understanding and Mitigating Over-Reliance in LLM Browsing Agents
- AAAI 2026 When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
2025
- NeurIPS 2025 EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
- EMNLP 2025 Can Large Language Models Identify Implicit Suicidal Ideation? An Empirical Evaluation
- EMNLP 2025 Understanding How Value Neurons Shape the Generation of Specified Values in LLMs
- TACL RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
- ACL 2025 Fraud-R1: A Multi-Round Benchmark for Assessing the Robustness of LLM Against Augmented Fraud and Phishing Inducements
- ACL 2025 Understanding the Repeat Curse in Large Language Models from a Feature Perspective
- ACL 2025 Rethinking Prompt-based Debiasing in Large Language Models
- COLING 2025 Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
2024
- Computational Linguistics A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
- NeurIPS 2024 D&B DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
- CoLM 2024 Model Autophagy Analysis to Explicate Self-consumption within Human-AI Interactions
Blog
- A Simple Guide and Best Practices for Using OpenClaw in Research2026/03
- Two New Concepts for Frontier AI Systems: “Context Engineering” and “Chain of Command”2025/10
- Misalignments and RL failure modes in the early stage of superintelligence2025/07ICLR Blogpost Track 2026 · also on LessWrong
Drafts and checkpoints
- Automatic Alignment Research — Part 1: Misbehavior Monitorabilitywork in progress, on Notion
- GitHub Issue-resolving agents: methods introductionwork in progress, on Notion
Services
- Area Chair: ACL Rolling Review (ARR)
- Conference reviewer: NeurIPS 2024, 2025; ICML 2025, 2026; ICLR 2025, 2026; AISTATS 2025, 2026; ARR (ACL, EMNLP, EACL); AAAI 2026; CVPR 2026; CoLM 2026
- Journal reviewer: Applied Artificial Intelligence; TALLIP; Transactions on Social Computing
- Workshop organizer: First Workshop on LLM Persona Modeling, NeurIPS 2025
- Teaching assistant: KAUST CS 325 Private Data Analysis, Fall 2025
Talks
- 2025.12 Who Is Controlling the Game? Opportunities and Challenges in Bidirectional Human–AI Alignment · EdinburghNLP, University of Edinburgh
Visiting and collaboration
- 2026.05 Visiting Student Researcher, Stanford Digital Economy Lab, Stanford HAI
- 2025.09 Visiting scholar, EdinburghNLP, University of Edinburgh
- 2025.04 Visiting scholar, THUNLP, Tsinghua University