Discover how one can construct custom LLM evaluators for specific real-world needs.
Considering the rapid advancements in the sector of LLM “chains”, “agents”, chatbots and other use cases of text-generative AI, evaluating the performance of language models is crucial for understanding their capabilities and limitations. Especially crucial to have the opportunity to adapt those metrics in line with the business goals.
While standard metrics like perplexity, BLEU scores and Sentence distance provide a general indication of model performance, based on my experience, they often underperform in capturing the nuances and specific requirements of real-world applications.
For instance, take a straightforward RAG QA application. When constructing a question-answering system, aspects of the so-called “RAG Triad” like context relevance, groundedness in facts, and language consistency between the query and response are vital as well. Standard metrics simply cannot capture these nuanced elements effectively.
That is where LLM-based “Blackbox” metrics turn out to be useful. While the thought can sound naive the concept behind LLM-based “blackbox” metrics is sort of compelling. These metrics utilise the facility of huge language models themselves to guage the standard and other elements of the generated text. By utilizing a pre-trained language model as a “judge”, we will assess the generated text in line with the language model’s understanding of the language and pre-defined criteria.
In this text, I’ll show the end-to-end example of constructing the prompt, running and tracking the evaluation.
Since LangChain is kinda of de-facto the preferred framework to construct chatbots and RAG, i’ll construct the appliance example on it. It would be easier to integrate into MVP and it has easy evaluation capabilities inside. Nonetheless, you should use every other frameworks you wish ot construct you own.
Major value of article — pipeline and prompts.
Let’s dive into the code and explore the strategy of creating custom evaluators. We’ll walk through just a few key examples and discuss their implementations.