Last year, we introduced Privatemode as a Confidential GenAI platform built on Confidential Computing, Contrast, and Confidential Containers to enable provider-excluded, end-to-end encrypted inference with open-source LLMs. Since then, we have operated Privatemode in production and gathered concrete learnings from real customer workloads.
In this follow-up talk, we move beyond architecture fundamentals and focus on the practical challenges of running confidential AI at scale. We present the dominant use cases we see today, including coding assistants, speech-to-text pipelines, and document processing, and derive the system requirements they impose.
We dive into real technical tradeoffs encountered in production. Topics include implementing prefix caching without cross-tenant leakage, evolving our transport security model to support enterprise TLS inspection without weakening confidentiality, and the current state of confidential GPU execution across NVIDIA Hopper and Blackwell.
Finally, we explore developer and UX challenges around attestation, including proxy-based, SDK-based, and browser-oriented approaches.
This session targets practitioners who want a realistic view of what it takes to deploy confidential AI in real-world environments, including what works, what breaks, and what still needs to evolve.