The model that performs best for adverse media screening may not be the best one for entity disambiguation, watchlist adjudication, or risk extraction. We continuously benchmark models by task because model selection affects both the quality of the work and the cost of providing it.