healthyml.org·há cerca de 1 mês When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment | Healthy ML
Yuxin Xiao is a Ph.D. candidate at MIT IDSS. His research focuses on building safe, robust, and trustworthy LLMs and advancing their reasoning and decision-making capabilities for healthcare and other