Agentic Infrastructure Risk: Securing Autonomous AI Systems in Production
About This Session
AI systems are rapidly evolving from passive models into autonomous agents capable of executing complex workflows across APIs, services, and cloud infrastructure. While this unlocks unprecedented automation, it also introduces a new and largely unaddressed category of risk.
In this talk, we explore agentic infrastructure risk - the failure modes, security vulnerabilities, and operational hazards that emerge when AI systems are allowed to take real-world actions. Drawing from experience building large-scale infrastructure at Meta, we examine how agent runtimes interact with APIs, tools, and micro-services, and why traditional security and reliability models fail in non-deterministic systems.
We will analyze critical emerging risks, including hallucinated actions, cascading retries, privilege escalation through tool misuse, and unbounded execution loops that can trigger system-wide incidents or cost explosions. Unlike traditional software failures, these risks are amplified by autonomy and feedback loops.
The session will introduce practical mitigation strategies:
* Guardrail architectures and action validation pipelines
* Policy-driven execution controls and zero-trust action models
* Observability for reasoning-driven, non-deterministic systems
* Cost containment and runtime safety mechanisms
Attendees will leave with a concrete, production-ready blueprint for designing AI systems that are secure, observable, and governed by design - enabling organizations to deploy autonomous agents without compromising safety or control.
Key Takeaways
- Identify new classes of AI risk introduced by autonomous agents
- Understand how agent failures differ from traditional system failures
- Learn to design guardrails and zero-trust execution layers
- Build observability for decision-making systems
Apply cost, safety, and governance controls in production
In this talk, we explore agentic infrastructure risk - the failure modes, security vulnerabilities, and operational hazards that emerge when AI systems are allowed to take real-world actions. Drawing from experience building large-scale infrastructure at Meta, we examine how agent runtimes interact with APIs, tools, and micro-services, and why traditional security and reliability models fail in non-deterministic systems.
We will analyze critical emerging risks, including hallucinated actions, cascading retries, privilege escalation through tool misuse, and unbounded execution loops that can trigger system-wide incidents or cost explosions. Unlike traditional software failures, these risks are amplified by autonomy and feedback loops.
The session will introduce practical mitigation strategies:
* Guardrail architectures and action validation pipelines
* Policy-driven execution controls and zero-trust action models
* Observability for reasoning-driven, non-deterministic systems
* Cost containment and runtime safety mechanisms
Attendees will leave with a concrete, production-ready blueprint for designing AI systems that are secure, observable, and governed by design - enabling organizations to deploy autonomous agents without compromising safety or control.
Key Takeaways
- Identify new classes of AI risk introduced by autonomous agents
- Understand how agent failures differ from traditional system failures
- Learn to design guardrails and zero-trust execution layers
- Build observability for decision-making systems
Apply cost, safety, and governance controls in production
Speaker
Nishant Gupta
Staff Software Engineer, Tech Lead - Meta (Meta SuperIntelligence Lab)
I am a Staff Software Engineer and Researcher at Meta, specializing in large-scale distributed systems and AI infrastructure. I build agentic systems where AI agents operate across APIs, services, and cloud environments with reliability and safety.
At Meta SuperIntelligence Lab, my work focuses on evaluation, guardrails, observability, and aligning autonomous systems with real-world outcomes. I also led elastic compute infrastructure managing ~30% of Meta’s capacity, driving significant efficiency gains. My research spans safe oversubscription systems and production AI infrastructure, with 90+ citations. I focus on building scalable, reliable systems that translate AI advances into real-world impact.
At Meta SuperIntelligence Lab, my work focuses on evaluation, guardrails, observability, and aligning autonomous systems with real-world outcomes. I also led elastic compute infrastructure managing ~30% of Meta’s capacity, driving significant efficiency gains. My research spans safe oversubscription systems and production AI infrastructure, with 90+ citations. I focus on building scalable, reliable systems that translate AI advances into real-world impact.