Chapter 8. AI-Driven Applications
In previous chapters, we demonstrated how to deploy model servers like vLLM on Kubernetes, package model data, and operate inference at scale. Building on that foundation, we will now shift from serving single models to architecting complete AI-driven applications where an LLM is just one of many components.
This chapter focuses on application architecture: how requests flow through a system, how context is retrieved or tools are invoked, and how state is maintained over time. We will introduce popular architectural patterns, the key components of AI application stacks, and the challenges of integrating LLMs into real-world applications. To maintain a clear focus on the architectural overview, we will keep the discussion at a high level. We will dive deeper into more concrete technical developments in the next chapter.
LLMs started their march of conquest into mainstream software as chatbots, with ChatGPT as their most prominent representative. Chat is still the dominant interaction pattern, but the software behind it has grown up. Modern AI apps wrap an LLM with application logic that fetches business context, calls internal systems, and writes state. The LLM inference service is a powerful component, but it does not reach into databases or call tools by itself.1
The application is in charge and uses the LLM for generation or reasoning. You will see where to use retrieval for grounding, when to orchestrate tool calls, and how to keep state across ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access