Notebooks
W
Weaviate
Transformation Agent Retrieval Benchmark

Transformation Agent Retrieval Benchmark

vector-searchagentsvector-databaseretrieval-augmented-generationllm-frameworksfunction-callingweaviate-recipesPythongenerative-aiweaviate-services

Open In Colab

Benchmarking Arctic 2.0 vs. Arctic 1.5 with Synthetic RAG Evals

Traditionally, AI system evaluation has relied heavily on human-written evaluation sets. Most notably, this process demands substantial time and resource investment, preventing most developers from creating and properly evaluating their AI systems.

The Weaviate Transformation Agent offers a breakthrough in AI evaluation by enabling the rapid generation of synthetic evaluation datasets!

In this notebook, we generate 2,100 synthetic questions for the Weaviate Blogs (1 for each) in 63 seconds!

We then use this dataset to report embedding recall (finding the source document in the haystack) between the Snowflake Arctic 2.0 and Arctic 1.5 embedding models. We find the following result:

ModelRecall@1Recall@5Recall@100
Arctic 1.50.79950.92450.9995
Arctic 2.00.84120.95460.9995

The Arctic 2.0 model demonstrates superior performance, with particularly notable improvements in Recall @ 1 (4.17% increase) and Recall @ 5 (3.01% increase).

Here is an overview of what this notebook illustrates:

Embedding Benchmark with the Transformation Agent

[2]
[4]
True

Weaviate Named Vectors

Create 2 HSNW indexes from the same property in Weaviate!!

On top of that, you can set different embedding models for each!

You can learn more about Weaviate Named Vectors here and the Weaviate Embedding Service here!

[48]

Simple Directory Parser to Load Weaviate's Blogs stored in Markdown Files into Memory

[5]

We have 2,160 blog chunks from Weaviate's Blog Posts!

[6]
2160

Here is an example of one of these blog chunks:

[7]
---
title: 'Accelerating Vector Search up to +40% with Intel’s latest Xeon CPU - Emerald Rapids'
slug: intel
authors: [zain, asdine, john]
date: 2024-03-26
image: ./img/hero.png
tags: ['engineering', 'research']
description: 'Boosting Weaviate using SIMD-AVX512, Loop Unrolling and Compiler Optimizations'
---

![HERO image](./img/hero.png)

**Overview of Key Sections:**
- [**Vector Distance Calculations**](#vector-distance-calculations) Different vector distance metrics popularly used in Weaviate. - [**Implementations of Distance Calculations in Weaviate**](#vector-distance-implementations) Improvements under the hood for implementation of Dot product and L2 distance metrics. - [**Intel’s 5th Gen Intel Xeon Processor, Emerald Rapids**](#enter-intel-emerald-rapids)  More on Intel's new 5th Gen Xeon processor. - [**Benchmarking Performance**](#lets-talk-numbers) Performance numbers on microbenchmarks along with simulated real-world usage scenarios. What’s the most important calculation a vector database needs to do over and over again?

Batch Import Blogs into Weaviate

[ ]
Inserted 500 blog chunks... (Time elapsed: 2.26 seconds)
Inserted 1000 blog chunks... (Time elapsed: 4.21 seconds)
Inserted 1500 blog chunks... (Time elapsed: 4.64 seconds)
Inserted 2000 blog chunks... (Time elapsed: 6.49 seconds)
[ ]
Attempting to retry failed objects...
Retry complete. Successfully uploaded 0 previously failed objects. 0 objects still failed.

Create Simluated User Queries with the Transformation Agent

Weaviate Users may be interested in knowing what embedding model they should use for their documents.

In order to know this, you will need some dataset of questions that you might expect your knowledge base to receive.

This can now be achieved at scale super quickly (2.1K questions in 63 seconds) with the Transformation Agent!

[ ]
{'workflow_id': 'TransformationWorkflow-e979d5f69b91575e7289ab61d559cafa', 'status': {'batch_count': 0, 'end_time': None, 'start_time': '2025-03-12 01:21:58', 'state': 'running', 'total_duration': None, 'total_items': 0}}
[ ]
{'workflow_id': 'TransformationWorkflow-e979d5f69b91575e7289ab61d559cafa',
, 'status': {'batch_count': 9,
,  'end_time': '2025-03-12 01:23:01',
,  'start_time': '2025-03-12 01:21:58',
,  'state': 'completed',
,  'total_duration': 63.097952,
,  'total_items': 2160}}
[62]
[TransformationResponse(operation_name='predicted_user_query', workflow_id='TransformationWorkflow-e979d5f69b91575e7289ab61d559cafa')]

Inspect Synthetic Queries produced from the Blog Content

[8]
[11]
Example #1
Blog Content:

---
title: 'Introducing the Weaviate Query Agent'
slug: query-agent
authors: [charles-pierse, tuana, alvin]
date: 2025-03-05
tags: ['concepts', 'agents', 'release']
image: ./img/hero.png
description: "Learn about the Query Agent, our new agentic search service that redefines how you interact with Weaviate’s database!"
---
![Introducing the Weaviate Query Agent](./img/hero.png)


We’re incredibly excited to announce that we’ve released a brand new service for our [Serverless Weaviate Cloud](https://weaviate.io/deployment/serverless) users (including free Sandbox users) to preview, currently in Alpha: the _**Weaviate Query Agent!**_ Ready to use now, this new feature provides a simple interface for users to ask complex multi-stage questions about your data in Weaviate, using powerful foundation LLMs. In this blog, learn more about what the Weaviate Query Agent is, and discover how you can build your own!

Let’s get started. :::note 
This blog comes with an accompanying [recipe](https://github.com/weaviate/recipes/tree/main/weaviate-services/agents/query-agent-get-started.ipynb) for those of you who’d like to get started. :::

## What is the Weaviate Query Agent

AI Agents are semi- or fully- autonomous systems that make use of LLMs as the brain of the operation. This allows you to build applications that are able to handle complex user queries that may need to access multiple data sources. And, over the past few years we’ve started to build such applications thanks to more and more powerful LLMs capable of function calling, frameworks that simplify the development process and more.
Predicted User Query:

How can I integrate the Weaviate Query Agent with my existing application to access multiple data sources using LLMs?

Example #2
Blog Content:

:::note 
To learn more about what AI Agents are, read our blog [”Agents Simplified: What we mean in the context of AI”](https://weaviate.io/blog/ai-agents). :::

**With the Query Agent, we aim to provide an agent that is inherently capable of handling complex queries over multiple Weaviate collections.** The agent understands the structure of all of your collections, so knows when to run searches, aggregations or even both at the same time for you. Often, AI agents are described as LLMs that have access to various tools (adding more to its capabilities), which are also able to make a plan, and reason about the response. Our Query Agent is an AI agent that is provided access to multiple Weaviate collections within a cluster. Depending on the user’s query, it will be able to decide which collection or collections to perform searches on.
Predicted User Query:

How does the Query Agent decide which Weaviate collections to perform searches on based on a user's query?

Example #3
Blog Content:

So, you can think of the Weaviate Query Agent as an AI agent that has tools in the form of Weaviate Collections. In addition to access to multiple collections, the Weaviate Query Agent also has access to two internal agentic search workflows:

-   Regular [semantic search](/blog/vector-search-explained) with optional filters
-   Aggregations

In essence, we’ve released a multi-agent system that can route queries to one or the other and synthesise a final answer for the user. ![Query Agent](img/query-agent.png)

### Routing to Search vs Aggregations

Not all queries are the same. While some may require us to do semantic search using embeddings over a dataset, other queries may require us to make [aggregations](https://weaviate.io/developers/weaviate/api/graphql/aggregate) (such as counting objects, calculating the average value of a property and so on). We can demonstrate the difference with a simple example.
Predicted User Query:

How do I decide when to use semantic search versus aggregations in the Weaviate Query Agent?

Example #4
Blog Content:

This allows you to provide the agent with instructions on how it should behave. For example, below we provide a system prompt which instructs the agent to always respond in the users language:

```python
multi_lingual_agent = QueryAgent(
    client=client, collections=["Ecommerce", "Brands"],
    system_prompt="You are a helpful assistant that always generated the final response in the users language."
    " You may have to translate the user query to perform searches. But you must always respond to the user in their own language."
)
```

## Summary

The Weaviate Query Agent represents a significant step forward in making vector databases more accessible and powerful. By combining the capabilities of LLMs with Weaviate's own search and aggregation features, we've created a tool that can handle complex queries across multiple collections while maintaining context and supporting multiple languages. The resulting agent can be used on its own, as well as within a larger agentic or multi-agent application.
Predicted User Query:

How do I integrate the Weaviate Query Agent with a larger agentic or multi-agent application?

Example #5
Blog Content:

Whether you're building applications that require semantic search, complex aggregations, or both, the Query Agent simplifies the development process while providing the flexibility to customize its behavior through system prompts. As we continue to develop and enhance this feature, we look forward to seeing how our community will leverage it to build even more powerful AI-driven applications. Ready to get started? Check out our [recipe](https://colab.research.google.com/github/weaviate/recipes/blob/main/weaviate-services/agents/query-agent-get-started.ipynb), join the discussion in our forum, and start building with the Weaviate Query Agent today!

import WhatsNext from '/_includes/what-next.mdx'

<WhatsNext />
Predicted User Query:

How can I customize the behavior of the Query Agent through system prompts?

The Generated Dataset can be found on HuggingFace here!

Benchmark Recall at 1, 5, and 100 (Arctic 1.5 vs. Arctic 2.0)

[ ]
Starting evaluation...
Processed 500 items...
  Current Arctic 1.5 - Recall@1: 0.8140, Recall@5: 0.9380
  Current Arctic 2.0 - Recall@1: 0.8780, Recall@5: 0.9760
Processed 1000 items...
  Current Arctic 1.5 - Recall@1: 0.8040, Recall@5: 0.9290
  Current Arctic 2.0 - Recall@1: 0.8630, Recall@5: 0.9660
Processed 1500 items...
  Current Arctic 1.5 - Recall@1: 0.8027, Recall@5: 0.9287
  Current Arctic 2.0 - Recall@1: 0.8493, Recall@5: 0.9567
Processed 2000 items...
  Current Arctic 1.5 - Recall@1: 0.7980, Recall@5: 0.9235
  Current Arctic 2.0 - Recall@1: 0.8390, Recall@5: 0.9540
Evaluation complete! Processed 2160 total items.
FINAL RESULTS:
Arctic 1.5 - Recall@1: 0.7995, Recall@5: 0.9245, Recall@100: 0.9995
Arctic 2.0 - Recall@1: 0.8412, Recall@5: 0.9546, Recall@100: 0.9995

Arctic Embed

Here are some more resources to learn about the Snowflake Arctic Embedding models!

Arctic Embed on GitHub
Arctic Embed 2.0 Research Report
Arctic Embed Research Report
Weaviate Podcast #110 with Luke Merrick, Puxuan Yu, and Charles Pierse!

Massive thanks to Luke and Puxuan for reviewing this notebook!

Arctic Embed on the Weaviate Podcast