Framework: AI Tool Evaluation Matrix
A structured decision matrix for evaluating AI tools before committing. Scores tools across seven weighted criteria to cut through marketing hype and make informed choices.
Use cases
Strategy & Planning, Operations & Workflow
Platforms
Model-Agnostic
Jump to a section
Workspace-ready resource
Copy the raw material, then check the setup, output contract, and failure modes.
# AI Tool Evaluation Matrix ## Instructions1. List the tools you are comparing across the top row.2. Score each tool from 1-5 on each criterion (1 = poor, 5 = excellent).3. Multiply each score by the weight.4. Sum the weighted scores per tool.5. Highest total score wins — but read the notes below before deciding. ## The Matrix | Criterion | Weight | Tool A | Tool B | Tool C ||-----------|--------|--------|--------|--------|| **Output Quality** | x3 | _/5 (=_) | _/5 (=_) | _/5 (=_) || Does it produce results good enough to use with minimal editing? Test with YOUR actual use cases, not the demo. | | | | || **Reliability** | x3 | _/5 (=_) | _/5 (=_) | _/5 (=_) || Does it produce consistent results? Does it fail gracefully or silently? Test the same input 3 times. | | | | || **Integration** | x2 | _/5 (=_) | _/5 (=_) | _/5 (=_) || How well does it connect to your existing workflow? API available? Webhooks? Works with your other tools? | | | | || **Speed to Value** | x2 | _/5 (=_) | _/5 (=_) | _/5 (=_) || How quickly can you go from sign-up to useful output? Hours? Days? Weeks of configuration? | | | | || **Cost Efficiency** | x2 | _/5 (=_) | _/5 (=_) | _/5 (=_) || What is the real cost? Include: subscription, API usage, time spent configuring, time saved. Net value, not sticker price. | | | | || **Data Handling** | x1 | _/5 (=_) | _/5 (=_) | _/5 (=_) || Where does your data go? Is it used for training? Can you export? What happens if you cancel? | | | | || **Longevity** | x1 | _/5 (=_) | _/5 (=_) | _/5 (=_) || Is this tool likely to exist and be supported in 12 months? Funded? Growing? Or burning runway? | | | | || **TOTAL** | **/70** | **__** | **__** | **__** | ## How to Score **5 — Excellent.** Best in class. No meaningful compromise.**4 — Good.** Works well with minor limitations you can live with.**3 — Adequate.** Gets the job done but you notice the gaps.**2 — Weak.** Requires significant workarounds.**1 — Poor.** Fails to meet basic expectations for this criterion. ## Evaluation Protocol ### Before scoring:1. Define your primary use case in one sentence. Score based on THIS use case, not the tool's entire feature set.2. Test each tool with the same 3 representative inputs from your real work.3. Time yourself: how long from "I want to do X" to "X is done acceptably"? ### Red flags (automatic disqualification regardless of score):- The tool requires your data but has no clear privacy policy.- The free tier is a bait-and-switch with aggressive upselling.- The core feature only works on one specific model or platform with no alternative.- The company has no visible team, funding, or track record. ### After scoring:- If two tools are within 5 points of each other, they are effectively tied. Choose based on gut feel about the team and trajectory.- If the winner has a score of 2 or below on any x3 criterion (Output Quality or Reliability), reconsider. A tool that is unreliable or produces poor output is not saved by being cheap and fast.- If no tool scores above 42/70 (60%), consider whether you need a tool at all, or whether a simpler approach (manual process, different model, custom solution) would serve better. ## Decision Log After completing the evaluation, record: - **Tools evaluated:** [list]- **Primary use case:** [one sentence]- **Winner:** [tool name]- **Key reason:** [one sentence — what tipped the decision]- **Reservations:** [what concerns remain]- **Review date:** [set a date 3 months out to reassess]Workspace translation
Turn this resource into an inspectable run.
Best next step
Use it as a review standard for the next output you save.
From resource to system
Use it once
Use the framework as a checklist, rubric, or decision aid.
Make it reusable
Attach it as project knowledge so future threads inherit the same criteria.
Decide after the run
Use it as a review standard for the next output you save.
Quality bar
Before using this, check the contract.
What input does this require?
What output should it produce?
Where can it fail?
What should a human review?
Recommended path
When to Use This
Use this before committing to any AI tool, platform, or service. The matrix is designed to cut through marketing hype and demo-driven excitement by forcing structured evaluation against criteria that actually matter for day-to-day use.
Especially useful when: your team is debating between competing tools, you are advising a client on AI tool selection, you are about to commit budget to an annual subscription, or you find yourself evaluating a new AI tool every week (the matrix forces discipline).
Why It Works
Weighted criteria reflect real-world importance. Output Quality and Reliability are weighted at x3 because they determine whether you actually use the tool after the first week. A tool that is fast to set up (x2) but produces unreliable output (x3) will score poorly, as it should. The weights encode the priority hierarchy that most people learn through painful experience.
The evaluation protocol prevents demo bias. Most AI tools are evaluated by watching a demo or running the tool's own example inputs. The protocol requires testing with YOUR actual use cases and YOUR real inputs. This is where most tools reveal their limitations.
The scoring anchors (1-5 with descriptions) prevent inflation. Without anchors, everyone scores everything 3 or 4. The explicit descriptions ("5 = best in class, no meaningful compromise") force honest assessment.
The red flag list provides automatic disqualification. Some issues are not matters of degree. No privacy policy is not a "score of 2 on Data Handling" — it is a reason not to use the tool at all. The red flags encode these deal-breakers.
The decision log creates accountability. Recording the reasoning and setting a review date prevents the common pattern of choosing a tool, forgetting why, and letting the subscription auto-renew long after the tool stopped being useful.
How to Customise
Adjust the weights. If integration is more important than output quality for your context (e.g. the tool connects critical systems), increase the Integration weight. The weights should reflect YOUR priorities.
Add criteria. Some teams need additional criteria: "Team adoption" (will people actually use it?), "Compliance" (does it meet regulatory requirements?), or "Customisability" (can you tailor it to your workflow?). Add rows with appropriate weights.
Create use-case-specific versions. An evaluation matrix for "AI writing tools" might add criteria like "Voice consistency" and "Format flexibility." An evaluation for "AI code tools" might add "Language support" and "IDE integration." Specialise the matrix for your most common evaluation scenarios.
Limitations
This framework is deliberately simple. It does not capture nuance that matters in some contexts: enterprise procurement requirements, security audit results, vendor lock-in analysis, or total cost of ownership modelling. For enterprise-scale decisions, use this matrix as a first pass to shortlist, then apply deeper due diligence.
The scoring is subjective. Two evaluators may score the same tool differently. For team decisions, have each evaluator score independently, then discuss and align where scores diverge by more than 1 point.
Model Notes
This framework is model-agnostic. It evaluates AI tools, not AI models. However, you can use this same structure to compare models (Claude vs GPT vs Gemini) by adjusting the criteria to: output quality, instruction following, context window, pricing, API reliability, and ecosystem.
Related Resources
Browse FrameworksSystem Prompt: Research Analyst
A system prompt for configuring an LLM as a structured research analyst that separates facts from interpretation, scores confidence, and flags gaps clearly.
Pattern / playbook seed
Research & Analysis · Strategy & Planning
Framework: Prompt Audit Checklist
A 15-point checklist for evaluating any prompt before putting it into production. Catches the most common prompt failures: vague instructions, missing constraints, absent error handling, and untested edge cases.
Knowledge / rubric seed
Operations & Workflow · Strategy & Planning
Need this operationalized?
Turn the pattern into a workspace system.
Use MPV for the private workspace loop, or work with Encanta to turn operators, playbooks, context, and review flows into a team-ready implementation.