AI Models Keep Changing. Due Diligence Technology Has to Change With Them

The model that performs best for adverse media screening may not be the best one for entity disambiguation, watchlist adjudication, or risk extraction. We continuously benchmark models by task because model selection affects both the quality of the work and the cost of providing it.

Threat.Digital model benchmarking process for selecting AI models for due diligence tasks

Threat.Digital uses large language models across a number of functions supporting due diligence firms, investigative researchers, third-party risk management platforms, compliance teams responsible for third-party risk, and AML enhanced due diligence teams.

When we started doing this, there was pretty much only one leading general-purpose model and there were few comparable alternatives. Model selection was relatively simple. Today, there are multiple major providers, several model families from each, different model sizes, and multiple reasoning levels that change how much time and compute a model uses before producing an answer.

AI model comparison chart showing intelligence, task time, and the Pareto front across multiple model providers
The model market now includes dozens of viable options across providers, model sizes, and reasoning settings. The Pareto front shows how quickly the tradeoffs between capability and performance become complicated. Source: Artificial Analysis

That is a much better market for building AI products. It is also a much more complicated one for any teams looking to leverage AI.

One model does not fit every due diligence or risk intelligence task

We use LLMs for a range of jobs, including:

  • Adverse media analysis
  • Risk extraction
  • Entity identification
  • Alias and AKA expansion
  • Sanctions and watchlist result adjudication
  • Risk-category classification
  • Content analysis and querying
  • Investigative workflows
  • Structuring information for downstream risk analysis

These tasks require different capabilities. Determining whether two companies with similar names are actually the same entity can require careful reasoning across names, locations, ownership information, and other identifying details. Extracting a defined set of risk categories from an article may be much more straightforward.

Using the most powerful model available for both tasks would be easy, but not necessarily better. The larger model may materially improve the difficult entity-resolution problem while producing no meaningful improvement on the classification task, but leading to a far greater cost for users.

So we evaluate the tasks separately.

Model selection is something that has to be maintained

The number of possible choices has increased quickly. A provider may offer several model tiers, with each one supporting multiple reasoning levels. Three model tiers with six reasoning settings already creates 18 possible configurations before comparing a second provider.

That does not mean we test every possible combination every day. It means we need a repeatable way to evaluate new options when they become relevant.

Threat.Digital has an internal model comparison system that lets us run defined due diligence use cases across different models and configurations. Depending on the task, we compare things such as accuracy, consistency, reasoning quality, latency, and cost.

Threat.Digital internal model comparison tool evaluating two AI models on an adverse media due diligence task
Internal model comparison arena. Basically Thunderdome for AI models - two models enter, one leaves!

Public AI benchmarks can be useful, but they do not tell us whether a model is better at resolving an ambiguous company match, identifying relevant adverse media, or reviewing a potential watchlist result. We need to test those questions using the work the model will actually perform.

Over time, that gives us something more useful than a list of model scores. It gives us experience with how different classes of models behave on due diligence and risk research tasks.

Better does not always mean bigger

One of the more useful things we have learned is that model performance is not a simple ladder.

A larger model can be better for a difficult reasoning task. A smaller model can perform just as well on a narrower task. In some cases, a smaller model given more reasoning time can outperform a larger model using a faster setting.

The goal is to find the sensible tradeoff between quality, speed, and cost for each application. In optimization, this is sometimes described as staying on the Pareto front: avoiding choices where another option is both better and cheaper.

For a TPRM or due diligence platform, this matters at scale. The cost difference between two models may look trivial for one request. It becomes significant when the same operation is applied across large volumes of companies, articles, names, watchlist records, and research results.

If a more expensive model materially improves the result, we want to use it. If a less expensive model produces the same result, paying more for the larger model serves no purpose.

A new model release is a test, not an automatic upgrade

The pace of model releases can make AI products appear to become outdated almost immediately. That is only true if the product is tightly tied to a particular model.

When a new model becomes available, we can evaluate it against the model already performing a given task. It may produce better results. It may be faster. It may produce equivalent results at a lower cost. Or it may offer no meaningful improvement at all.

The important part is having the ability to tell the difference.

For a company considering AI in a TPRM, due diligence, or enhanced due diligence program, this is worth asking vendors about. The question is not simply which AI model they use today. It is how they determine which models should perform different functions, how they test new releases, and how they know when a change actually improves the product.

AI model selection is not just a decision made when the system is built. It is part of maintaining the system.

Our users should not have to follow every model release or understand the differences between providers, model sizes, and reasoning settings. That is work we do behind the product.


Learn more about DiligenAI, Threat.Digital's AI-powered due diligence, monitoring, and risk intelligence platform.