AI Production: From Agents to Evaluation

Alps Wang

Alps Wang

Sep 25, 2026 · 1 views

The Evolving AI Systems Engineering Landscape

The QCon AI New York 2026 article highlights a critical shift in AI development: AI engineering is becoming synonymous with systems engineering. This evolution is driven by the increasing autonomy and complexity of AI systems, moving the focus from individual model behavior to the robust operation of the entire system. Key themes emerging from the conference sessions include the intricate challenges of agent authorization and identity management, where traditional human-centric models are inadequate for software agents acting as users. The need for bounded execution authority, secure delegated authority across agent chains, and auditable multi-hop tool calls are paramount. This necessitates a re-evaluation of security and access control mechanisms, moving beyond simple role-based access to more dynamic and context-aware authorization frameworks.

Furthermore, the article underscores the growing importance of operationalizing AI at scale. Sessions on guardrails for operations agents, like LinkedIn's Kubernetes Ops Agent, demonstrate the necessity of robust engineering controls, including rate limits, access controls, and peer approval, to prevent unintended consequences when AI agents interact with production infrastructure. The discussion around shared inference platforms, exemplified by Netflix's approach to scaling ML and GenAI, emphasizes the architectural and organizational trade-offs involved in consolidating model serving. This consolidation aims to improve efficiency and allow ML practitioners to focus more on modeling, but it introduces complexities in managing diverse latency requirements and defining clear boundaries between business and model logic. The economic aspect, particularly inference costs, latency, and token usage, is now a first-class architectural constraint, demanding careful optimization.

Finally, the article addresses the crucial aspect of post-deployment AI evaluation. DoorDash's Alchemy platform illustrates the challenges of trusting and teaching AI moderation systems where ground truth is often ambiguous or non-deterministic. The focus on building evaluation harnesses that support shadow-mode testing, backtesting, and metrics tied to incident reduction rather than just model accuracy is a significant takeaway. The ability for AI judgments to become training data for lower-cost models represents a promising avenue for cost optimization and continuous improvement. Overall, these discussions collectively paint a picture of AI engineering maturing into a discipline that deeply integrates principles from distributed systems, security, platform engineering, and SRE, with the ultimate goal of building reliable, observable, controllable, and economical AI systems for production.

Key Points

  • AI engineering is increasingly becoming systems engineering, focusing on system behavior over model behavior.
  • Identity and authorization for autonomous agents are critical, requiring new models beyond traditional human-centric approaches.
  • Robust guardrails and engineering controls are essential for AI agents interacting with production infrastructure.
  • Shared inference platforms offer efficiency gains but introduce architectural trade-offs in managing diverse model requirements.
  • Post-deployment AI evaluation, including handling ambiguous ground truth and using AI judgments for retraining, is crucial for production AI systems.

Article Image


📖 Source: From Agent Authorization to AI Production Evaluation: QCon AI New York 2026

Related Articles

Comments (0)

No comments yet. Be the first to comment!