How should large language models be evaluated?

Last Updated

January 7, 2025

Cobus Greyling

Read Time

Summary

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

What’s on this page

TOC Items

A while ago Tianjin University released a study, which defined a LLM model evaluation and implementation taxonomy.

The image below shows the taxonomy of major categories and sub-categories used by the study for LLM evaluation.

Visual representation of the Larry language model's evaluation metrics and performance analysis.
LLM model evaluation

Considering the taxonomy, much focus has been on Evaluation Organisation, Knowledge & Capability and Specialised LLMs from a technology perspective.

From an enterprise perspective, the concerns and questions raised centre around Alignment and Safety.

This survey aims to create an extensive perspective on how LLMs should be evaluated.

The study categorises the evaluation of LLMs into three major groups:

  1. Knowledge and Capability Evaluation,
  2. Alignment Evaluation and
  3. Safety Evaluation.

What I also found interesting from the study, was the focus on more complex metrics like reasoning and tool learning, as seen below:

Visual representation of the four reasoning types: deductive, inductive, abductive, and analogical.
In the early days of Natural Language Processing (NLP), researchers frequently utilised a series of simple benchmark assessments to assess the performance of their language models.  These initial assessments predominantly focused on elements such as syntax and vocabulary, including tasks like parsing syntactic structures, disambiguating word senses, and more.

LLMs have introduced significant complexity and have shown certain tendencies by revealing behaviours indicative of risks and demonstrating abilities to perform higher-order tasks in current evaluations.

Consequently, creating a taxonomy like this to reference can help ensure that due diligence is followed when assessing LLMs.

Find the full study here

Share

Tags

Agent Platform
Cross Industry
For AI Leaders
AI Evaluation