PDF Files
Copyright 2025 Google LLC.
Setup
Configure your API key
To run the following cell, your API key must be stored in a Colab Secret named GOOGLE_API_KEY. If you don't already have an API key, or you're not sure how to create a Colab Secret, see Authentication for an example.
Download and inspect the PDF
Install the PDF processing tools. You don't need these to use the API, it's just used to display a screenshot of a page.
Reading package lists... Done Building dependency tree... Done Reading state information... Done The following NEW packages will be installed: poppler-utils 0 upgraded, 1 newly installed, 0 to remove and 35 not upgraded. Need to get 186 kB of archives. After this operation, 697 kB of additional disk space will be used. Get:1 http://archive.ubuntu.com/ubuntu jammy-updates/main amd64 poppler-utils amd64 22.02.0-2ubuntu0.8 [186 kB] Fetched 186 kB in 0s (722 kB/s) Selecting previously unselected package poppler-utils. (Reading database ... 126111 files and directories currently installed.) Preparing to unpack .../poppler-utils_22.02.0-2ubuntu0.8_amd64.deb ... Unpacking poppler-utils (22.02.0-2ubuntu0.8) ... Setting up poppler-utils (22.02.0-2ubuntu0.8) ... Processing triggers for man-db (2.10.2-1) ...
This PDF page is an article titled Smoothly editing material properties of objects with text-to-image models and synthetic data available on the Google Research Blog.
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0
100 6538k 100 6538k 0 0 29.6M 0 --:--:-- --:--:-- --:--:-- 29.6M
Look at one of the pages:
page-image-1.jpg sample_data test.pdf
Upload the file to the API
Start by uploading the PDF using the File API.
Try it out
Now select the model you want to use in this guide, either by selecting one in the list or writing it down. Keep in mind that some models, like the 2.5 ones are thinking models and thus take slightly more time to respond (cf. thinking notebook for more details and in particular learn how to switch the thiking off).
The pages of the PDF file are each passed to the model as a screenshot of the page plus the text extracted by OCR.
1560
In addition, take a look at how the Gemini model responds when you ask questions about the images within the PDF.
If you observe the area of the header of the article, you can see that the model captures what is happening.
Learning more
The File API lets you upload a variety of multimodal MIME types, including images, audio, and video formats. The File API handles inputs that can be used to generate content with model.generateContent or model.streamGenerateContent.
The File API accepts files under 2GB in size and can store up to 20GB of files per project. Files last for 2 days and cannot be downloaded from the API.
- Learn more about prompting with media files in the docs, including the supported formats and maximum length.
- Learn more about to extract structured outputs from PDFs in the Structured outputs on invoices and forms example.