Gemini AI

How to Use Google Gemini API with Python for Beginners

Introduction to Google Gemini API

Google's Gemini models are a family of highly capable, multimodal AI models designed to understand text, images, audio, and code. This tutorial will guide you through using the Gemini API in Python, from setting up your first script to handling advanced multimodal requests.

Prerequisites and Setup

Before coding, you need an API key and the official Python SDK.

  1. Go to Google AI Studio and generate an API key.
  2. Ensure you have Python 3.9 or higher installed.
  3. Install the generative AI SDK via pip:
pip install -q -U google-generativeai

Basic Configuration

Start by configuring the SDK with your API key. It's best practice to use environment variables rather than hardcoding the key.

import google.generativeai as genai
import os

genai.configure(api_key=os.environ["GEMINI_API_KEY"])

Generating Text (Gemini Pro)

The standard model for text tasks is Gemini Pro. Here's how to generate a simple response.

model = genai.GenerativeModel('gemini-pro')
response = model.generate_content("Explain quantum computing to a 5-year-old.")
print(response.text)

Streaming Responses

For longer responses, you might want to stream the text as it's being generated, which improves perceived latency in applications.

response = model.generate_content("Write a long story about a space explorer.", stream=True)
for chunk in response:
    print(chunk.text, end="")

Multimodal Inputs: Text and Images (Gemini Pro Vision)

Gemini Pro Vision can understand images alongside text prompts. First, ensure you have the Pillow library installed (pip install Pillow).

import PIL.Image

img = PIL.Image.open('example_image.jpg')
model = genai.GenerativeModel('gemini-pro-vision')

response = model.generate_content(["Describe what you see in this image:", img])
print(response.text)

Configuring Safety Settings

The Gemini API includes built-in safety filters for categories like harassment, hate speech, and sexually explicit content. You can adjust these thresholds.

safety_settings = [
  {
    "category": "HARM_CATEGORY_HARASSMENT",
    "threshold": "BLOCK_ONLY_HIGH"
  },
  {
    "category": "HARM_CATEGORY_HATE_SPEECH",
    "threshold": "BLOCK_NONE" 
  }
]

response = model.generate_content(
    "Your prompt here",
    safety_settings=safety_settings
)

Warning: Adjusting safety settings to lower thresholds means your application may output offensive content. Use with caution.

Error Handling

Robust applications should gracefully handle API errors, such as rate limits or blocked content.

try:
    response = model.generate_content("Some prompt")
    if response.prompt_feedback.block_reason:
        print("Prompt was blocked:", response.prompt_feedback)
    else:
        print(response.text)
except Exception as e:
    print(f"An error occurred: {e}")

Conclusion

The Google Gemini API offers a straightforward yet incredibly powerful way to integrate cutting-edge AI into your Python applications. By mastering text generation, streaming, and multimodal inputs, you can build sophisticated tools ranging from chatbots to automated image analyzers.