Agent Shortlist

Author

Lucas Powell

Lucas Powell

Founder, Growth 8020 · Editor, Agent Shortlist

Lucas Powell runs Growth 8020, an AI-first B2B marketing studio that builds and operates AI agent workflows for companies who want senior-level marketing output without the senior-level headcount. The work spans cold outreach automation, lead research, content production, and customer-support deflection — across Claude, GPT, Gemini, Lindy, Relevance AI, n8n, Paperclip, and roughly two dozen other platforms tested in production.

He started Agent Shortlist after watching too many teams — including his own — burn through AI tools that promised value and didn't deliver. The publication is built on a simple premise: builders deserve honest reviews from people who've actually used the tools, not affiliate-padded listicles or vendor-sponsored content disguised as editorial.

Every platform on the site has been tested in real workflows, not vendor demos. Pricing is verified daily against vendor sources and published as an open dataset on GitHub under CC-BY-4.0 — so other builders can cross-check the numbers and use them in their own tools. Verdicts are opinionated and come with reasoning. When something fails, we say so, by name.

The other half of the time, Lucas is heads-down on Growth 8020 client work — building the same agent workflows we review, in real production environments with real budgets and real consequences. The two feed each other: the publication's recommendations are stress-tested by real client deployments, and the client work surfaces the failure modes that most reviews never mention.

Areas of expertise

  • AI agent platform evaluation and selection
  • Cold outreach and lead-research agent design
  • Model routing and token-cost optimisation
  • No-code and low-code agent orchestration (Lindy, n8n, Relevance AI, Paperclip)
  • Coding agent stacks (Claude Code, Cursor, Cline, Aider)
  • Voice AI deployment patterns (Retell, Vapi, Bland)

Platforms tested in production (2026)

  • Claude Code, Cursor, Cline, Windsurf, Aider, Augment Code, Amp, OpenAI Codex, GitHub Copilot
  • OpenClaw, Hermes, Paperclip, Manus AI
  • Lindy, Relevance AI, Stack AI
  • Retell AI, Vapi, Bland AI, ElevenLabs Conversational
  • n8n, Make
  • Anthropic Claude API, OpenAI API, Google Gemini, DeepSeek, Moonshot Kimi, z.ai GLM

Editorial standards

Three rules. One: every platform we review has been tested in a real workflow, not in a vendor demo. If we haven't shipped real work on it, we don't review it. Two: affiliate relationships are disclosed on the about page and never affect verdicts or rankings. We've turned down sponsorship offers that came with editorial strings. Three: when something fails, we say so. Naming the failure mode is more useful than burying it. Builders making real decisions deserve real information, not "Lindy is good — try it for 14 days."

Where to find Lucas

Articles by Lucas

Jul 3, 2026 · 7 min read

Augment Code vs Cursor vs Claude Code (2026): Which Wins for Your Team

Augment Code vs Cursor vs Claude Code compared for 2026 — pricing, codebase size, agent autonomy, and the decision rule for picking the right coding agent.

Jul 3, 2026 · 8 min read

Claude Extra Usage in 2026: What It Costs, When It's Worth It, and the Three Billing Pools People Confuse

Claude extra usage is billed at API rates on Pro, Max 5x, and Max 20x. Real per-model math, three billing pools people confuse, and when upgrading beats topping up.

Jul 3, 2026 · 5 min read

How to Check Claude Code Usage (2026): Every Method That Actually Works

Four ways to check Claude Code usage: the /status and /cost commands, transcript files under ~/.claude, and a free dashboard that rolls up all your history.

Jun 30, 2026 · 8 min read

Does Claude Work in Microsoft Copilot Studio? (2026)

Does Claude work in Microsoft Copilot Studio in 2026? The supported model options, how Claude integrates via Azure, and the practical workarounds when it doesn't.

Jun 23, 2026 · 9 min read

Claude Code vs OpenAI Codex (2026): Which Coding Agent Wins

Claude Code or OpenAI Codex in 2026? Pricing, ChatGPT bundling, agent autonomy, and the tasks where each wins after side-by-side testing.

Jun 23, 2026 · 8 min read

Claude Opus vs Sonnet 4.6 (2026): When Each Model Actually Wins

Claude Opus 4.8 costs 1.6x more than Sonnet 4.6 but isn't always 1.6x better. The 70/20/10 routing pattern and cost math at production volume.

Jun 22, 2026 · 8 min read

Claude Pro vs Max vs API: When Each One Actually Wins

Claude Pro at $20, Max at $100, or pay-per-token API? Cost math at three usage tiers, crossover points, and the hybrid pattern serious users run.

Jun 22, 2026 · 5 min read

Hermes vs Aider: when each one is the right pick

Hermes and Aider are both open-source and model-agnostic but answer different questions. The decision rule, cost math, and when each wins.

Jun 22, 2026 · 5 min read

Hermes vs Cline: which open-source agent fits your workflow

Hermes and Cline are open-source and model-agnostic but built for different work. The honest decision rule, cost math, and when each is the right pick.

Jun 22, 2026 · 4 min read

Hermes vs OpenHands: the open-source agent comparison that matters

Hermes and OpenHands are both top-tier open-source agents — but they're built for different jobs. Decision rule, cost math, and when each is the right pick.

Jun 22, 2026 · 9 min read

How to Avoid Hitting Claude Usage Limits (2026 Builder's Guide)

Patterns that let builders use Claude 3-5x more without hitting limits. Prompt caching, model routing, subagents, hooks, and when to use the API.

Jun 17, 2026 · 7 min read

Evals as PRDs: How AI Teams Are Replacing Specs With Tests

Evals are becoming the new PRD for AI agents — quantifiable tests that define 'done' for coding agents. Frameworks, patterns, failure modes.

Jun 17, 2026 · 10 min read

How to Create an AI Agent: A Tested Builder's Guide (2026)

How to create an AI agent in 2026: four paths from no-code to fully custom, with our platform pick, time to first agent, and real cost.

Jun 17, 2026 · 9 min read

Loop Engineering: How to Design Self-Prompting AI Agents

Loop engineering — four patterns that turn AI agents into autonomous systems. Heartbeats, crons, hooks, goals, and platforms that ship each.

Jun 17, 2026 · 8 min read

Multi-Agent AI: When to Use It, When to Skip It, What Actually Works

Multi-agent AI compared honestly, the three patterns that work in production, the four that don't, and the cost math that decides which is right.

Jun 11, 2026 · 10 min read

The best AI coding agents in 2026

The best AI coding agents in 2026 — Claude Code for agentic work, Cursor for IDE-native, Cline for open-source. Honest verdicts on 12 tools, no affiliate padding.

Jun 11, 2026 · 8 min read

The best AI voice agents in 2026

Best AI voice agents in 2026 — Retell for production, Vapi for developer flexibility, Bland for turnkey outbound, ElevenLabs for premium voice quality.

Jun 11, 2026 · 8 min read

The best no-code AI agent builders in 2026

Best no-code AI agent builders in 2026 — Lindy for fastest ship, Relevance AI for complex outreach, Stack AI for document Q&A, Manus for autonomous research.

May 17, 2026 · 6 min read

AI Agent Task Selection: The ARR Framework (Autonomous, Recurring, Reviewable)

AI agent task selection made simple. The ARR framework — Autonomous, Recurring, Reviewable — decides which tasks belong with an AI agent and which don't.

May 17, 2026 · 8 min read

Director vs doer: the mindset shift that separates working AI agents from broken ones

Stop prompting. Start directing. The mindset change builders need to make once they move from chatbots to agents, and the practices that come with it.

May 17, 2026 · 5 min read

Hermes vs Cursor: a comparison nobody else makes, and why it matters

Hermes and Cursor get compared by people who don't know they're different categories. What each does, why the question matters, and which to pick.

May 17, 2026 · 12 min read

How much does it cost to build an AI agent in 2026?

AI agent development costs in 2026: no-code ($30–$300/mo), low-code ($50–$300/mo), custom builds ($2k–$50k first month). Cost-per-task, hidden items.

May 17, 2026 · 8 min read

The lethal trifecta: the AI agent security trap nobody warns you about

Three capabilities safe alone, catastrophic combined: private data, internet access, untrusted input. How the AI agent security trap works and how to break it.

May 3, 2026 · 7 min read

Self-hosted AI is bigger than you think

Three of the top productivity AI tools by usage are self-hosted and open source. That's not the narrative. Here's the usage data behind it.

May 1, 2026 · 6 min read

Why two open-source agents quietly own the productivity category in 2026

Two open-source agents own 95% of productivity tokens on public model-marketplace leaderboards. Why the market concentrated this fast.

Apr 29, 2026 · 6 min read

OpenClaw vs Hermes: which open-source agent should you self-host in 2026?

OpenClaw vs Hermes head to head, the two open-source agents builders actually run. Trade-offs that matter and the usage data behind the choice.

Apr 27, 2026 · 8 min read

Skills vs MCP servers vs subagents: the architectural map for builders

Five concepts in Claude Code overlap with each other. Most explainers stop at definitions. Here's when to actually use which.

Apr 24, 2026 · 9 min read

What Claude Skills actually are (and why most people are getting them wrong)

Most builders think Claude Skills are saved prompts. The architecture is different, and it's the reason Skills are actually useful.

Apr 22, 2026 · 9 min read

The zero-human company: five roles AI agents are quietly replacing

A new orchestration pattern: entire org charts populated by AI personas, with humans as the board. Five roles where it's working, and where the math breaks.

Apr 20, 2026 · 9 min read

Claude Code vs Cursor (2026): Real Differences After 6 Months of Daily Use

Claude Code or Cursor in 2026? Pricing math at three usage tiers, the IDE-fork vs CLI tradeoff, model lock-in, and the specific tasks where each one wins.

Apr 17, 2026 · 9 min read

The 2026 AI Agent Shortlist: 8 Platforms Worth Your Team's Time

Tested eight AI agent platforms on real builder workflows. These are the ones that survived. Ranked, verdicted, and ready to shortlist.

Apr 15, 2026 · 8 min read

The real cost of Claude at scale in 2026

Per-token math on real Claude workloads — support agents, customer-deflection at 50k tickets/month, prompt caching. Five cost levers ranked by impact.

Apr 13, 2026 · 11 min read

Where AI agents actually deliver ROI in 2026 (and where the math doesn't work)

Five patterns where AI agents pay for themselves, and the vendor math you should ignore. With concrete numbers from real production deployments.

Apr 9, 2026 · 5 min read

How to Start Using AI Agents in Your Business (Without Breaking Anything)

The non-technical owner's guide to AI agents: the mindset shift, the task audit trick, and why one boring automated task beats a robot empire on day one.

Apr 6, 2026 · 10 min read

AI Agents for Finance: Where They're Actually Working in 2026

Five finance workflows where AI agents are delivering measurable ROI, with real numbers. Plus where the math still doesn't work and which platforms to use.

Apr 1, 2026 · 11 min read

The 5 most common AI agent use cases (and which platform fits each)

What builders are using AI agents for in 2026: five patterns: customer support, sales, research, code, ops. With the platform we'd pick for each.

Mar 30, 2026 · 7 min read

Email AI Agents: The Best Tools for Inbox Automation in 2026

The email AI agents worth using in 2026 — inbox triage, reply drafting, follow-up sequences, and outreach. Tools, costs, and what actually works.

Mar 25, 2026 · 5 min read

AI Agent Model Routing: Cut Your API Bill by 60% Without Losing Quality

Brain-and-muscle model routing: use expensive models for planning, cheap models for execution. Real cost breakdowns and the routing logic that makes it work.

Mar 20, 2026 · 13 min read

AI Agent Observability: What to Monitor and How

Best AI agent observability tools and what to instrument. LangSmith, Paperclip, OpenTelemetry, Datadog, the four metrics every production agent needs.

Mar 17, 2026 · 13 min read

AI Agent Guardrails: How to Not Delete Your Database in 9 Seconds

Seven AI agent guardrails every production deployment needs — approval gates, action boundaries, budget caps, blast radius. Why 9 seconds ended Pocket OS.

Mar 12, 2026 · 15 min read

Best AI Agent Orchestration Platforms in 2026: Ranked

The 8 best AI agent orchestration platforms in 2026 — Paperclip, n8n, LangGraph, CrewAI, Lindy, AutoGen, Azure, Vertex. Ranked with decision rules.

Mar 9, 2026 · 9 min read

AI Agent Workflow Design: Patterns That Ship in Production

AI agent workflow design, the eight patterns that ship in production, the three that don't, and the decision tree for picking the right one for your job.

Mar 4, 2026 · 11 min read

The best AI agent frameworks in 2026: LangGraph, CrewAI, AutoGen, and what to pick

The best AI agent frameworks compared — LangGraph, CrewAI, AutoGen, Semantic Kernel. Which fits which workflow, and when to skip them entirely.

Mar 2, 2026 · 10 min read

AI Agent Skills and Memory: How to Make Agents Get Better Over Time

Skills files, context management, and routines turn a one-trick agent into a system that improves. The architecture behind agents that compound over time.