AI tool intelligence

Find the right AI products for the work.

Our live catalog contains 2,444 researched AI products. Each record keeps current product positioning, fit, limitations, pricing signals, and official destinations together for a clearer decision.

Search products, capabilities, and workflowsBrowse the index ↓
2,444AI products in the catalog
DatedAvailability and evidence checks
IndependentFit, limitations, and alternatives

Public AI tool catalog

Search this public selection by the work you need to improve.

Search the complete catalog by product, capability, workflow, category, or tag. Every result opens a profile built from the same live record.

CE
Catalog evaluation

MLOps

Censius

Censius is an AI observability platform for machine learning teams monitoring production models. It brings together model details, input data, logs, and monitoring insights to identify performance, drift, data quality, and bias issues, then guide root cause analysis and remediation.

Model MonitoringModel ExplainabilityModel DriftRoot Cause AnalysisPython SDK
Review
EA
Catalog evaluation

Developer tools

Evidently AI

Evidently AI is an open-source framework for evaluating, testing, and monitoring LLM applications, RAG systems, AI agents, and predictive machine-learning models. Teams can use built-in or custom metrics, generate test cases, run evaluations locally or in the cloud, and track quality over time.

LLM evaluationML monitoringData driftRAG evaluationOpen source
Review
FA
Catalog evaluation

AI observability

Fiddler AI

Fiddler AI is an enterprise AI control plane for monitoring, evaluating, securing, and governing agents, LLM applications, and traditional ML models. It gives teams execution context and decision lineage, applies guardrails, and supports continuous evaluation throughout the AI lifecycle.

Agentic observabilityLLM monitoringGuardrailsAI governanceModel evaluation
Review
HE
Catalog evaluation

Developer tools

Helicone

Helicone provides an AI gateway and observability platform for teams building LLM applications. Developers can route requests across providers, track usage and costs, debug sessions, manage prompts and datasets, and configure reliability features such as caching, rate limits, and automatic fallbacks.

LLM ObservabilityAI GatewayCost TrackingPrompt ManagementModel Routing
Review
HU
Catalog evaluation

AI development tools

Humanloop

Humanloop was acquired by Anthropic, and its platform was sunset on September 8, 2025. Before the sunset, it provided enterprise tooling for LLM evaluation, prompt management, and production observability, helping technical and non-technical teams collaborate on AI product development.

Prompt managementLLM evaluationsObservabilityModel monitoringDeveloper SDKs
Review
LA
Catalog evaluation

Developer tools

Latitude

Latitude is an open-source platform for tracing production AI agents, identifying recurring failures, and monitoring fixes. Teams can capture agent telemetry, search sessions, turn failure patterns into signals, generate evaluations, and dispatch coding agents with trace context.

OpenTelemetryAgent monitoringFailure analysisSelf-hostingMCP
Review
PR
Catalog evaluation

AI development

PromptLayer

PromptLayer is an AI engineering platform for managing prompt versions, evaluating LLM and agent behavior, and tracing production activity. Teams can let domain experts edit prompts visually while engineers connect releases, requests, costs, latency, and quality signals to the relevant prompt versions.

Prompt managementLLM evaluationsAgent observabilityPrompt versioningSDKs
Review
RA
Discontinued record

MLOps platforms

Radicalbit

Radicalbit provides enterprise AI infrastructure for deploying and serving ML, computer-vision, and LLM applications. Its products cover real-time ML pipelines, model observability, open-source AI monitoring, and an AI Gateway for routing, guardrails, usage controls, and visibility across AI traffic.

Model monitoringLLM observabilityData drift detectionAI GatewayEvent stream processing
Review

How the index works

A catalog built for decisions, not rankings.

Commercial relationships can fund the research. They never change product status, fit, limitations, evidence, or the alternatives we show.

Verify what exists

We check availability, product direction, pricing signals, and when the supporting evidence was last reviewed.

Evaluate operational fit

We assess the work it supports, integration requirements, ownership, risk, and where human judgment remains necessary.

Keep the comparison honest

Commercial relationships are disclosed, while credible limitations and alternatives stay visible.

Evaluating a product for a live workflow?

Bring the decision to our engineers