Building AI Apps with Google AI Studio and Gemini
Published by O'Reilly Media, Inc.
Harnessing Gemini’s multimodal capabilities with Gemini and Google AI Studio
Course outcomes
- Understand Gemini 3.0 architecture and its capabilities
- Develop applications utilizing Gemini’s crossmodal reasoning
- Optimize prompts for multimodal data processing
- Leverage Google AI Studio to implement the Gemini API for practical applications, from chatting with PDFs to interactively querying videos
Join expert Lucas Soares to dive into in-depth exploration of building applications using the Gemini API, focusing on the latest Gemini 3.0 models, both flash and advanced. You’ll gain a comprehensive understanding of Gemini’s architecture, its multimodal capabilities, and best practices for prompt design and model integration. You’ll get hands-on experience developing end-to-end applications, leveraging Gemini’s state-of-the-art performance in areas like reasoning, coding, and multimedia processing.
What you’ll learn and how you can apply it
- Learn the fundamentals of Gemini and its API
- Learn how to prototype applications with Google AI Studio
- Explore the advanced features of Gemini models for app development
- Design effective prompts for enhanced model performance
- Integrate Gemini models into real-world applications
- Improve multimodal application accuracy and efficiency
This live event is for you because...
- You’re a software developer, AI engineer, or data scientist interested in building cutting-edge applications with multimodal AI capabilities.
Prerequisites
- A basic understanding of modern LLMs
- Experience with programming languages such as Python
- Familiarity with RESTful APIs and web services
Recommended preparation:
- Explore Gemini Shortcuts (expert playlist)
- Read Gemini’s Technical Report (documentation)
Recommended follow-up:
- Take Using LLMs for Software Engineering (live online course with Chelsea Troy)
- Read Hands-On Large Language Models (book)
- Read GenAI on Google Cloud (book)
Schedule
The time frames are only estimates and may vary according to how the class is progressing.
Introduction (20 minutes) Presentation and demonstration: Introduction to Google’s Gemini; Gemini models and capabilities; Gemini app; Google AI Studio Hands-on exercise: Set up environment and API keys Break
First steps of building with Gemini (60 minutes) Presentation: Walkthrough of Google AI Studio for prototyping Gemini apps; building a Gemini app Hands-on exercises: Talk, show, and share with the multimodal live API using Gemini 3.0; build a simple multimodal chatbot with Gemini Flash Q&A Break
Developing a process for building Gemini apps (70 minutes) Presentation: Designing and testing prompts for text, image, and audio inputs; prototyping, building, and deploying apps Hands-on exercises: Develop multimodal prompts with Google AI Studio; build a multimodal chat with Docs app Q&A Break
Systematically improving your Gemini apps (30 minutes) Presentation: Testing your application and assessing latency, cost, and quality in the Google AI Studio Dashboard Hands-on exercise: Improve your app systematically across latency, cost, and quality
Agentic workflows with Gemini (60 minutes) Presentation: Introduction to agents and agentic workflows Hands-on exercise: Build a simple personal assistant Q&A
Your Instructor
Lucas Soares
Lucas Soares is a machine learning engineer who has worked at K1 Digital and Biometrid, where he developed computer vision and NLP models for applications such as document verification, OCR-based applications, and recommender systems. Lucas has also developed various ML models, including neural networks, Siamese networks, convolutional neural networks, LSTMs, and genetic algorithms.