My Take on the AI VC Benchmark · A Blogger's Diary

I Made 11 AI Agents Do My Job. Here’s What Happened.

A personal dive into the AIMultiple VC benchmark — and why I’m not panicking (yet).
JD
Jamie Dalton 19 July 2026 · 7 min read

Let me paint you a picture. It’s a Tuesday morning, I’ve got three cups of coffee in me, and I’m staring at a spreadsheet of 200+ startups. My job? Figure out which ones are actually worth a partner’s time. It’s the kind of grunt work that makes you question your life choices — but also the kind that, if you’re good, makes you indispensable.

Then I read the AIMultiple AI VC Benchmark. Somebody went and turned my Tuesday mornings into a test for AI agents. Eleven of them. Two tasks. And the results? They’re weird, they’re messy, and they tell a much more interesting story than “AI is coming for your job.”

“Strength in one venture capital workflow did not carry over to the other.” — The report’s quiet bombshell.

That line stopped me cold. Because if you’ve been following the AI hype, you’d think these models are universally brilliant. But the benchmark shows a different reality: the leader for deal sourcing was not the leader for competitor mapping. It’s like hiring a brilliant chef and then being surprised they can’t fix your car.

Let’s talk about the numbers that made me sweat

The benchmark split the work into two classic VC analyst workflows:

  • Deal sourcing: Find young AI startups (founded in 2025–2026) with under $5M funding, founded by Turkish diaspora, headquartered outside Türkiye. Sounds specific? That’s the point.
  • Competitor mapping: Map all the competitors for a target company (in this case, the research firm AIMultiple itself). A messy, judgment-heavy task.

Here’s the visual that slapped me in the face:

AgentVC opportunity + CDD avg (success %)
Sonnet 558.5%
Fable 552.7%
Gemini 3.1 Pro44.3%
GPT-5.6 Terra42.0%
GPT-5.6 Sol41.8%
Opus 4.834.1%
GPT-5.6 Luna27.7%
GLM 5.227.2%
MiniMax M326.8%
Gemini 3.5 Flash20.9%
Kimi K30%

Look at that spread. From 58.5% down to 0%. And the top performer? Not even a passing grade in most schools. This isn’t “AI is superhuman.” This is “AI is a very fast intern who sometimes hallucinates funding rounds.”

Deal sourcing: the great middle

Fable 5
72%
Gemini 3.1 Pro
67.3%
GPT-5.6 Terra
63.6%
Sonnet 5
57.5%
GPT-5.6 Sol
57.3%
Opus 4.8
55.9%
GPT-5.6 Luna
55.4%
MiniMax M3
48.9%
Gemini 3.5 Flash
41.1%
GLM 5.2
37.1%

Most agents clustered in the 55–72% range for sourcing. The benchmark authors note that “the qualifying companies were founded within the previous 18 months, which puts them past the reach of training data.” In other words, the AI can’t just memorize the answer. It has to search, verify, and reason. And it’s pretty good at that — just not great.

🔍
Image: My mental model of an AI agent sourcing deals
It’s like watching a very fast intern who forgets to check their sources.

Competitor mapping: the real nightmare

Sonnet 5
47.9%
Fable 5
45%
GPT-5.6 Sol
31.3%
GPT-5.6 Terra
19.9%
GLM 5.2
17.4%
Gemini 3.1 Pro
16.6%
Kimi K3
15.2%
Opus 4.8
12.3%
MiniMax M3
4.8%
Gemini 3.5 Flash
0.6%

This is where the wheels came off. Only two agents scored above 45%. The rest? A graveyard of single-digit scores. Why? Because competitor mapping is judgment, not just verification. It asks: “Who is a real competitor?” — which is a question even humans argue about.

The report puts it perfectly: “sourcing rewards mechanical verification… competitor mapping demands a judgment about where a niche research firm’s market ends.” And that judgment, my friends, is the secret sauce. It’s what separates a junior analyst from a partner.

🧠
Image: My brain trying to define a “competitor”
It’s nuanced, messy, and apparently very hard for AI.

What I learned about AI (and about myself)

Reading this benchmark felt like looking in a mirror. The tasks are exactly what I do. And the AI’s struggles are exactly my struggles: verifying founder origin, deciding if a company is “close enough” to be a competitor, not hallucinating funding data.

But here’s the thing — the benchmark doesn’t show AI replacing me. It shows AI augmenting me. The top-performing agents can handle the grunt work: scanning, filtering, pulling data. But they still need a human to make the final call. To decide if that Turkish-sounding surname is actually evidence. To judge if that SEO tool is a competitor or just a neighbor.

“Every figure needs a source the agent opened. Bare homepages and search-result pages do not count.” — The rule that separates good research from bullshit.

I also deeply appreciated the leak control the authors implemented. They discovered that if the scoring rubric was accessible to the model, the model would read it and cheat. So they isolated it. That’s the kind of rigor that makes me trust the results. It’s also a sobering reminder: AI will take shortcuts if you let it. Just like a lazy analyst.

So, will AI take my job?

Not yet. Not for the high-judgment stuff. But will it make my job different? Absolutely. I can already see myself using a tool like Fable 5 or Sonnet 5 to do the first-pass sourcing, then spending my time on the deeper due diligence. That’s not a layoff — that’s a promotion.

The benchmark’s final line stuck with me: “teams selecting an agent should test it on the workflow they plan to automate rather than relying on a single leaderboard.” In other words: don’t trust the hype. Test. Validate. And remember that a model that’s great at sourcing might be terrible at competitor mapping.

For me, that’s the real takeaway. AI isn’t a monolith. It’s a toolbox. And we’re the ones who get to choose the right tool for the job.

— Jamie Dalton, recovering spreadsheet jockey.

📄 Original Article (AIMultiple)

The following is the full text of the original AIMultiple benchmark as published. I’ve kept it intact because the data deserves to be read in full.

AI VC Benchmark: 11 AI Agents on Real Venture Capital Tasks

Ezgi Arslan, PhD. with Berk Kalelioğlu · updated on 18 Jul 2026

Agents AI Matériel d'IA Applications GenAI L'IA dans l'industrie Fondements de l'IA RAG Modèles d'IA

We converted two AI VC analyst workflows, AIMultiple runs commercially, deal sourcing, and competitor mapping, into benchmarks with human-verified ground truth and scored 11 AI agents on them. The leader changed between the two tasks: strength in one venture capital workflow did not carry over to the other.

See the results, task designs, and the scoring method, including the leak control that keeps models from inflating their own scores by reading the rubric.

Venture capital benchmark results

VC opportunity + CDD average — success rate (%)

ModelScore
Gemini 3.5 Flash20.9%
MiniMax M326.8%
GLM 5.227.2%
GPT-5.6 Luna27.7%
Opus 4.834.1%
GPT-5.6 Sol41.8%
GPT-5.6 Terra42.0%
Gemini 3.1 Pro44.3%
Fable 552.7%
Sonnet 558.5%

Each of the 11 models ran each task once. Scores are out of 100. Kimi K3 produced no scorable deal-sourcing run and is recorded as 0.

Both sit inside the sector AIMultiple itself comes from: the sourcing run hunts for young companies building on artificial intelligence, and the due diligence target earns its revenue by researching and benchmarking AI products. Every question, therefore, lands in the fastest-moving corner of venture capital, where new companies form faster than training data updates, which is exactly where an investor needs fresh research rather than recalled knowledge.

Diaspora deal sourcing results

VC opportunity identification — success rate (%)

AgentScore
GLM 5.237.1%
Gemini 3.5 Flash41.1%
MiniMax M348.9%
GPT-5.6 Luna55.4%
Opus 4.855.9%
GPT-5.6 Sol57.3%
Sonnet 557.5%
GPT-5.6 Terra63.6%
Gemini 3.1 Pro67.3%
Fable 572.0%

Most of the field cleared half the points and landed in a broad middle band, with a wide gap between the leader and the tail. The rubric explains the shape. A row earns points when every criterion carries evidence, so an agent scores at the pace it can verify founder origin, founding year, and funding, not at the pace it can list plausible names. The qualifying companies were founded within the previous 18 months, which puts them past the reach of training data, so the spread reflects search depth and verification discipline rather than memorized knowledge. Even the leading agent left a substantial share of the points on the table, meaning full recall of a fresh, evidence-gated gold set remains out of reach for current agents working under a time budget.

Competitor mapping results

Commercial due diligence (competitors) — success rate (%)

AgentScore
GPT-5.6 Luna0.0%
Gemini 3.5 Flash0.6%
MiniMax M34.8%
Opus 4.812.3%
Kimi K315.2%
Gemini 3.1 Pro16.6%
GLM 5.217.4%
GPT-5.6 Terra19.9%
GPT-5.6 Sol31.3%
Fable 545.0%
Sonnet 547.9%

The distribution is split into three tiers, each mapped to a scoring mechanism. The top pair covered several of the target’s competitor categories and caught part of the close-peer group. The middle band is what the rubric produces when an agent maps the well-known analyst and review names cleanly and misses the small, recent peers nearest to the target: solid execution, penalized twice through lost recall points and flat miss penalties. The near-zero tail is the most instructive, because a score that low does not require an empty file. Penalties for missed peers and for returned non-competitors can cancel everything a tidy CSV earns, so a confident, well-formatted map of the wrong companies ends up worth the same as no map at all.

Why results of the two tasks diverge

Sourcing rewards mechanical verification of five stated criteria; competitor mapping demands a judgment about where a niche research firm’s market ends, with no criteria list to check against. That judgment separated the field far more than formatting or speed did, and it is the reason the leader changed between columns while several agents swapped positions sharply. Strength in one research workflow says little about the next one, so teams selecting an agent should test it on the workflow they plan to automate rather than relying on a single leaderboard.

Venture capital benchmark methodology

Each task ships as a spec file we authored, which fixes the analyst role, the output file, the tool access, and the time budget before any run starts. The table summarizes those specs:

TaskAnalyst workTime limit
Diaspora deal sourcingFilling an early-stage fund's pipeline: scanning the market for startups that fit the investment thesis and documenting each candidate for the partners120 min
Competitor mappingSupporting an investment decision: mapping the market around a deal target so the investor sees who it competes with before committing capital90 min

The time limit caps each run: an agent must research, verify, and deliver its CSV within that window, and deal sourcing gets the longer budget because it verifies dozens of candidate companies against five criteria while competitor mapping investigates a single target.

Both place the agent in a role a junior analyst holds at a venture capital fund, and both require a CSV with atomic cells, resolving source URLs, and no commentary. All facts are judged as of a frozen snapshot date, and gold sets are re-verified quarterly because funding data drifts.

Task 1: Diaspora deal sourcing

The agent acts as a deal-sourcing analyst at an early-stage fund focused on Turkish diaspora founders. It must return AI companies that meet all of these investment criteria as of the snapshot date:

  • Funding below $5M. Total disclosed capital raised, not a valuation or revenue figure. If funding is undisclosed, the company qualifies if positive evidence shows the amount falls below the threshold. Otherwise, it must be excluded.
  • Founded in 2025 or 2026. Anything earlier fails.
  • Founder origin with public evidence. At least one founder was born in Türkiye, educated at a Turkish institution, or born to Turkish parents. A Turkish-sounding surname is not evidence.
  • Headquartered abroad. The company itself sits outside Türkiye.
  • AI-focused. Infrastructure-layer and application-layer companies both qualify.

Each row pairs every figure with a source URL the agent opened. A fabricated or unreachable URL invalidates the row it supports.

How it is scored

Scoring runs against a hidden ground truth: a gold set of verified qualifying companies the model never sees. Four dimensions apply:

  • Recall against a verified gold set of qualifying companies.
  • Precision against a trap set of near-misses, each failing exactly one criterion. Named disqualifiers carry fixed penalties.
  • Data accuracy on a fixed random sample of rows drawn with a recorded seed, so a scoring pass can be reproduced: headquarters, founding year, funding within tolerance, and sources that support the claim beside them.
  • Format compliance: column order, atomic cells, numeric funding, valid CSV.

Padding a list with plausible-but-unverified companies loses more in precision than it gains in recall. The design rewards fewer, fully verified rows.

Task 2: Competitor mapping

The agent conducts commercial due diligence on a target company for a potential investor. For the scored run, we pointed the agents at our own company: the target is AIMultiple itself, which let us verify the ground truth against first-hand knowledge of who competes with us. The task: identify the target’s direct competitors, evidence each with public data, and return them as a CSV document.

The scope rules define what counts as a competitor:

  • Substitutability. A company qualifies when it offers a substitutable product or service to overlapping customers, in any category where the target competes.
  • Full category coverage. The target operates across multiple categories. The agent must cover competitors in every one of them, not the most obvious one alone.
  • Exclusions. Companies the target writes about, partners with, or lists, but does not compete with. Adjacent players in a neighboring category. The target itself.

A thorough map for a target of this breadth runs 8 to 15 competitors. Each row includes the competitor’s LinkedIn page, total disclosed funding (or Undisclosed), headquarters country, founding year, and LinkedIn headcount, with a source URL beside each metric. Two reasoning columns sit alongside the metrics: a differentiator versus the target and a rationale for inclusion, each required to be specific rather than generic. A funding figure must be backed by a funding source.

How it is scored

Scoring runs against a hidden ground truth, organized by competitor category. Five dimensions apply:

  • Precision. Every returned non-competitor reduces the score. Named disqualifiers, such as SEO tools and the AI model vendors, that the target merely benchmarks carry fixed penalties. A soft anti-padding guard reduces excess weight when a single category accounts for most of the found set. Competitors carry two weight levels. A small set of must-not-miss peers, the firms closest to the target’s core work, weigh three times a standard entry. A genuine competitor absent from the ground truth still earns credit after expert review, so an incomplete gold set does not punish a correct find.
  • Graduated from penalties. Independent of recall weight, each must-not-miss peer the agent fails to return subtracts a flat penalty sized by closeness. Missing the target’s nearest peers is treated as a core due diligence failure, not a routine recall gap.
  • Funding accuracy. Sampled rows checked within tolerance, with Undisclosed and N/A (public) accepted as correct answers where they apply.
  • Metric evidence and reasoning. Required metrics present, atomic, sourced, and accurate within tolerance. The differentiator and rationale cells are scored for specificity against boilerplate.

Why the criteria resist pattern matching

Both tasks are built to block the shortcuts agents take under time pressure:

  • Guessing fails. Where a figure is not public, the scored answer is an evidenced gap, not an estimate. An invented number costs points twice, once on accuracy and once on precision, so admitting the unknown outperforms filling it.
  • The answers sit past the first search page. Both gold sets lean on young or thinly documented companies that training data has not caught up with, so scores depend on live, layered search rather than recall.
  • Evidence beats inference. A plausible signal, a familiar-sounding name or a high-ranking page, does not establish a fact; a citable record does. Both rubrics score the source next to the claim, not the claim by itself.
  • Completeness and precision pull against each other. Padding with unverified names costs more than it earns, while an overly cautious short list loses recall. The scoring pushes agents toward the balance a fund expects from an analyst: every finding verified, no genuine finding skipped.

Building the tasks

Each benchmark began as a real workflow within our business development and research operations. We converted original analyst specs into task files with a fixed structure: role, eligibility rules, exact output columns, tool access, and a time limit. Rubrics live in separate files with per-criterion points: format and arithmetic run through automated checks, and judgment calls go to a human reviewer.

Three design rules carried across both tasks:

  • Every figure needs a source the agent opened. Bare homepages and search-result pages do not count.
  • Undisclosed beats invented. Writing Undisclosed for a private pre-seed round scores; inventing a number costs points twice, once for accuracy and once for precision.
  • Atomic output. One value per cell, exact column order. Structure violations cost points before content is judged.

Leak control

Early runs taught us the most important lesson in methodology. When the rubric file was reachable from the model’s working directory, the model read it and copied the ground truth. Scores rose for reasons unrelated to research ability.

Every rubric now carries a harness-isolation requirement: the model receives the prompt file, and nothing else, and the evaluator applies the rubric afterward. Any run where the rubric was reachable is treated as invalid. Benchmark builders who skip this control will overstate what their agents can do.

What the rubrics measure

The rubrics share a scoring philosophy across both tasks:

  • Recall against verified ground truth. Each qualifying company or competitor in the gold set was confirmed by a human researcher against public records before any model was scored.
  • Precision with named traps. Both ground truth sets carry near-misses that fail exactly one criterion and are named disqualifiers with fixed penalties, so pattern matching without verification gets caught.
  • Negative points for fabrication. A hallucinated funding figure or a dead source URL subtracts points instead of scoring zero. Honest gaps beat confident errors.
  • Drift management. Funding, valuations, and headcounts change. Each rubric carries a snapshot date, marks drift-prone data, and requires re-verification before every scoring pass.
  • Mixed checking. Recall, precision, and format checks run automatically. Founder-origin evidence and inclusion rationales go to a human reviewer.

Low scores come from making the task genuinely hard, not from crushing good answers with caps. Recall, precision, accuracy, and honesty each pull on a different failure mode, so a model cannot compensate for weak research with confident formatting.

Why venture capital work tests AI agents well

Venture capital firms run on analyst labor. Deal flow screening and due diligence are multi-step research tasks. Each one requires web research, source verification, and structured output. Each one also produces plausible-looking wrong answers, which makes them hard to fake and useful to score.

These properties make the VC analyst work a strong benchmark domain:

  • Verifiable ground truth. Funding rounds, founding years, and headcounts can be checked against public records.
  • Fabrication pressure. Models that cannot find a figure tend to invent one. The tasks expose this directly.
  • Real economic value. A fund that automates parts of due diligence changes its own cost structure. The benchmark measures work someone pays for today.

Further readings

  • AIM Agentic Marketing Benchmark
  • AI-Based Stock Trading: Which Gen AI Tool Is Better
  • Agentic AI Finance Benchmark

FAQ

Do these benchmarks show AI replacing venture capital analysts?
No. They show that current agents can handle specific, well-defined research tasks with moderate success, but they struggle with judgment-heavy work like competitor mapping.

Why do the tasks target private companies?
Private companies have less public data, which forces agents to rely on live search and verification rather than memorized training data.

Can a fund reuse these tasks internally?
Yes, the methodology and task designs are open and can be adapted to a fund’s specific investment thesis.


Citation: Ezgi Arslan, PhD. and Berk Kalelioğlu (2026) – "AI VC Benchmark: 11 AI Agents on Real Venture Capital Tasks". Published online at AIMultiple.com. Accessed 18 July 2026.

Ezgi Arslan, PhD.
Analyste du secteur
Berk Kalelioğlu
Chercheur en IA

No comments:

Post a Comment