
Education
Demystifying Safety in Generative AI: From Alignment to Policy Enforcement
Ends:
Information Sciences Building
TBD
LERSAIS/DINS/DOCTORAL GUILD SEMINAR
Demystifying Safety in Generative AI: From Alignment to Policy Enforcement
Nathalie Baracaldo, PhD
Senior Research Scientist and Master Inventor, IBM Research
Alumna, IS PhD Program
As agentic AI systems proliferate and large language models become ubiquitous, ensuring AI safety has never been more critical, and at the same time, more complex. Safeguarding generative AI requires addressing risks across multiple layers: data curation, training methodology, model harness design, and deployment infrastructure.
This talk examines key aspects of the AI training process, focusing on alignment techniques and unlearning mechanisms, then turns to a critical challenge: policy enforcement in enterprise GenAI systems. I will highlight discrepancies in what policy means to different communities and some of the very different ways policies are enforced in each silo. After that, I will explain common pitfalls that practitioners encounter and offer concrete recommendations for building safer systems.
Nathalie Baracaldo is a Senior Research Scientist and Master Inventor at IBM Research in San Jose, California, where her work focuses on building trustworthy AI systems. She has extensive experience delivering impactful machine learning solutions that are highly accurate, withstand adversarial attacks, and protect data privacy. Her current research focuses on safeguarding generative AI through unlearning and alignment techniques. She served as the principal investigator for the DARPA GARD program, leading efforts to extend and maintain the Adversarial Robustness Toolbox (ART) for red teaming evaluations. She also led IBM's federated learning initiative and co-edited two books: "Federated Learning: A Comprehensive Overview of Methods and Applications" (Springer, 2022) and "Machine Unlearning for Governance of Foundation Models" (2026). Her research has been published in top AI and security conferences, earning multiple best paper awards and thous
Sources: pitt_events
