Fact-check note: Revised September 4, 2026. The original 2025 directory is not a current catalog; product availability, features, prices, terms, and model versions change rapidly.

A long tool list is not a technical evaluation. Start with a defined task, users, affected people, baseline, alternatives, prohibited uses, risk tolerance, integration constraints, and accountable owner. Create a dated candidate inventory from first-party documentation, then test the complete system in your environment.

Use a repeatable evidence card

DimensionEvidence to record
IdentityProduct, provider, edition, region, model/version, review date, retirement policy
Task qualityRepresentative held-out results, error severity, citations, accessibility, human-review burden
Safety/securityAbuse tests, prompt injection, data leakage, tool permissions, incident history and response
Privacy/dataPurpose, training use, retention, deletion, subprocessors, regions, export, access
OperationsLatency, availability, quotas, logs, support, fallback, rollback, change notices
EconomicsUsage, seats, storage, integration, review, monitoring, incidents, migration and exit
GovernanceOwner, approvals, legal review, prohibited use, monitoring and re-evaluation triggers

Run a controlled comparison

  1. Freeze candidates, configurations, prompts, tools, and data.
  2. Use routine, rare, adversarial, multilingual, accessibility, privacy, and high-consequence cases.
  3. Prespecify quality, safety, latency, cost, and review thresholds.
  4. Repeat stochastic runs and retain inputs, outputs, tool traces, errors, and versions.
  5. Pilot with bounded permissions, fallback, stop criteria, and an accountable go/no-go decision.

Verify claims and terms

Confirm features, supported platforms, model access, pricing, licensing, and data terms in dated first-party sources. “Free,” “private,” “enterprise-ready,” “open source,” and “compliant” require exact scope. Provider assurance does not establish that a customer configuration or use is safe or lawful.

Plan for change and exit

Test data export, workflow portability, dependency removal, model replacement, contract termination, and record retention. Monitor quality and harms after deployment and re-evaluate model, policy, price, provider, integration, population, or regulation changes. Do not silently refresh a ranking while retaining an old publication context.

Compare underlying systems with AI model evaluation, define measures using LLM evaluation metrics, and govern selection through AI governance best practices.