Blog

Pushing a RAG prototype to production

| Gert Jan Spriensma

Crafting the Benchmark Answer Set

One of the first things we set up was a Benchmark Answer Set. We initially believed that having access to a large Q&A database would make this easy. It turned out that choosing questions that covered the full range of nuances was hard. Identifying and matching the right sources was also tricky, particularly when different sources discuss similar topics. So we opted for an iterative approach in which improvements were made progressively, together with domain experts and their feedback. This way we didn’t have to wait until the Benchmark Answer Set was ready; we could already start optimising the solution with the limited content we had.

Building the Evaluation Framework

Setting up an evaluation framework itself is relatively straightforward. Fine-tuning the metrics so they reflected our goals took quite a bit more effort. Metrics like recall look simple at first glance, but it is easy to get the details wrong, and we had to keep re-checking our methods to make sure they still measured what we wanted. Although our focus was on directional trends rather than absolute figures, having precise measurements made the insights a lot more useful. The main metrics we used were Recall, Precision/Mean Reciprocal Rank (MRR), and an LLM-based correctness metric, which together gave a good picture of how the system performed.

System Components: From Prototype to Beta

Going from a prototype to a beta version means integrating and refining multiple system components. If you have a prototype, you probably have the following components already set up:

• A basic pipeline for cleaning, chunking, and embedding

• A retrieval system using similarity searches

• A mechanism for generating answers

As we moved to the beta phase, the system grew to include more layers:

• A more elaborate pipeline for data processing

• A classification layer for smarter routing of queries

• A query expansion layer to break down and broaden the scope of questions

• A reranker to combine and prioritise results

• A guardrails layer to validate answers and sources

Project Attitude: Detail-Oriented and Hands-On

Throughout the project, our success depended less on the metrics themselves and more on our willingness to examine every part of the pipeline. That meant going through the details, such as unanswered questions and overlooked sources, and constantly testing and tweaking small parts of the solution. Domain expertise helps, but what mattered more was working with the data, recognising patterns, and making informed adjustments to the many components that can be tuned.

Enhancing Semantic Match

The most important task is to improve the semantic match between a question and the content from which the answer can be extracted. In our project, we plotted the embeddings in 2D and could clearly see the difference between the most important content (blue) and questions (red). Fortunately for us, there was already content in place that acted as a bridge between the questions and the main content. If you don’t have that bridge, you can consider strategies like HyDE or synthetic content generation to fill the gap.

Scatter plot of embeddings projected to 2D, with content and query points forming separate clusters

The 2D projection of our content (blue) and query (red) embeddings.

With query expansion, you can focus on different elements of the question (for example, a specific element of your business) to touch on multiple types of content in your database. The problem with this approach is that you need to bring the results from different queries together, merge, and deduplicate them. As you are now getting a lot more results, a reranker layer becomes important.

A reranker is a slower but more capable model (compared to an embedding model) that measures the similarity between the question and the content. In our case, reranking proved very difficult, as commercial offerings such as Cohere and Jina didn’t work for us, leaving us with larger and slower models. The reason was the very specific domain we operated in and Dutch as the main language. For more general applications, rerankers would probably have worked well.

This part of the project shows that you cannot ‘just’ apply a best practice and expect great results; you have to test and figure out what works for your specific case. We expected that implementing reranking would be fairly trivial, and in the end most of the development time went into this layer, without a perfect result.

And sometimes you need some luck

We got a big boost from the timely release of OpenAI’s new embedding models, which improved the similarity matching and lifted our recall by 10%-15%. That jump came from nothing more than switching from text-embedding-ada-002 to text-embedding-3-large. We also fine-tuned an embeddings model specific to our domain, but unfortunately the results were far from perfect, probably due to the language.

Stay tuned for the next post, where I’ll cover the second part of this two-month project, setting up guardrails, fine-tuning prompts and preparing for the beta launch.

Gert Jan Spriensma

Author

Gert Jan Spriensma

Experienced AI engineer who builds production-ready AI systems.

LinkedIn

Keep up with the latest

Sign up for our newsletter and get our views on the latest in data & AI.

Let's talk about your data.

Get in touch with our team at contact@mozaik.ai or use the form below.

Or visit us at our office: Pakhuis De Hoop, Breestraat 59, Amersfoort.

Pakhuis De Hoop, Breestraat 59, Amersfoort

We only use your details to reply to your message. Read more in our privacy policy.