AI tool intelligence

Find the right AI products for the work.

Our live catalog contains 2,444 researched AI products. Each record keeps current product positioning, fit, limitations, pricing signals, and official destinations together for a clearer decision.

Search products, capabilities, and workflowsBrowse the index ↓
2,444AI products in the catalog
DatedAvailability and evidence checks
IndependentFit, limitations, and alternatives

Public AI tool catalog

Search this public selection by the work you need to improve.

Search the complete catalog by product, capability, workflow, category, or tag. Every result opens a profile built from the same live record.

EA
Catalog evaluation

Developer tools

Evidently AI

Evidently AI is an open-source framework for evaluating, testing, and monitoring LLM applications, RAG systems, AI agents, and predictive machine-learning models. Teams can use built-in or custom metrics, generate test cases, run evaluations locally or in the cloud, and track quality over time.

LLM evaluationML monitoringData driftRAG evaluationOpen source
Review
HU
Catalog evaluation

AI development tools

Humanloop

Humanloop was acquired by Anthropic, and its platform was sunset on September 8, 2025. Before the sunset, it provided enterprise tooling for LLM evaluation, prompt management, and production observability, helping technical and non-technical teams collaborate on AI product development.

Prompt managementLLM evaluationsObservabilityModel monitoringDeveloper SDKs
Review
PR
Catalog evaluation

AI development

PromptLayer

PromptLayer is an AI engineering platform for managing prompt versions, evaluating LLM and agent behavior, and tracing production activity. Teams can let domain experts edit prompts visually while engineers connect releases, requests, costs, latency, and quality signals to the relevant prompt versions.

Prompt managementLLM evaluationsAgent observabilityPrompt versioningSDKs
Review
AN
Discontinued record

Machine learning

Anote

Anote is a human-centered AI platform for creating labeled datasets, improving domain-specific models, and evaluating AI systems. It combines annotation workflows and human feedback with fine-tuning, benchmarks, synthetic data, and developer tools for teams building reliable AI applications.

Data annotationLLM evaluationFine-tuningHuman feedbackSynthetic data
Review

How the index works

A catalog built for decisions, not rankings.

Commercial relationships can fund the research. They never change product status, fit, limitations, evidence, or the alternatives we show.

Verify what exists

We check availability, product direction, pricing signals, and when the supporting evidence was last reviewed.

Evaluate operational fit

We assess the work it supports, integration requirements, ownership, risk, and where human judgment remains necessary.

Keep the comparison honest

Commercial relationships are disclosed, while credible limitations and alternatives stay visible.

Evaluating a product for a live workflow?

Bring the decision to our engineers