ShipSift

Formerly IdeaMan · 7-pass multi-model research engine

Research that argues with itself, so your roadmap doesn't have to.

ShipSift runs your question through a 7-pass, multi-model research pipeline. It gathers cited evidence from the web and podcasts, labels every finding by persona and stance, puts it in front of a jury of competing AI models, and hands you a verdict you can defend.

Models and sources ShipSift orchestrates

One question. Models and sources from:

Anthropic
OpenAI
Mistral
xAI
Moonshot
DeepSeek
Perplexity
Firecrawl
Vercel
Anthropic
OpenAI
Mistral
xAI
Moonshot
DeepSeek
Perplexity
Firecrawl
Vercel
Anthropic
OpenAI
Mistral
xAI
Moonshot
DeepSeek
Perplexity
Firecrawl
Vercel
Anthropic
OpenAI
Mistral
xAI
Moonshot
DeepSeek
Perplexity
Firecrawl
Vercel

Trademarks belong to their owners. ShipSift is not affiliated with or endorsed by them.

How it fits together

One question in, one decision out

Ask about a company, a persona, an idea or a feature. ShipSift fans out across deep web research, targeted search, full-page reads and podcast transcripts, then converges on one memo.

Web research
Search
Page reads
Podcasts
ShipSift
Findings
Jury
Memo
Questions to ask

The 7-pass workflow

Not a chatbot: a pipeline of separate stages, each logged with its own health report.

01Research

A deep web research agent finds what's actually been published: prices, customers, case studies, job postings. Figures missing from snippets get re-checked on the full page.

  • perplexity-agent

02Evidence

Targeted searches, page reads and podcast episodes, with a quote pulled for every claim. A claim whose quote isn't in the source is dropped.

  • search
  • page reads
  • podcasts

03Triage

A classifier labels each finding by relevance, persona, job and stance, and flags any figure its source doesn't contain.

  • relevance
  • persona
  • job
  • stance

04Jury

Three different AI models react in character as each of your buyer personas.

  • 3 models
  • 3 labs

05Red team

A separate model builds the strongest case against the idea.

  • counter-case

06Judge

A frontier model writes the memo: a verdict per persona, a scorecard, and the questions to ask real people.

  • go / reshape / park

07Cross-check

An independent classifier re-scores the jury. Disagreements are shown, not hidden.

  • independent

Live run log

Every run leaves a receipt

This is a replay of a real run: what ran, what it kept, what it dropped and why.

$ shipsift run --persona "Merchandiser · Mid-market" --question "Would $5M+ Shopify apparel brands pay for sourced competitor intel?"
▸ research      perplexity-agent · 3 pages re-checked          ✓ 1m 37s
▸ evidence      5 queries → 5 results · 1 page read            ✓
▸ alexandria    3 queries → 30 podcast results
                asked: "competitor price monitoring software for apparel brands" …
▸ podcasts      30 pieces → 44 claims · 41 kept · 3 dropped (quote not found)
▸ extraction    57 claims · 3 dropped (5%) · 54 kept           ✓ 1m 1s
▸ triage        relevance · persona · job · stance · 1 figure flagged   ✓
▸ jury          mistral · gpt · grok × persona  3/3            ✓
▸ red team      counter-case written                            ✓
▸ judge         memo · verdict: RESHAPE · number check clean    ✓
▸ cross-check   15/15 re-scored                                 ✓
run complete · 104 calls · 0 failed · 5m 31s

What you can research

One evidence base, four kinds of question

Research companies, buyer personas, product ideas and app features side by side, against the same evidence.

Companies

Who they sell to, what they charge, who uses them, sourced.

  • › customers · pricing · case studies
  • › every line carries its quote

Personas

How merchandisers, marketers, designers or any role actually work, and what they'd pay for.

  • › job-to-be-done · budget line
  • › reactions labeled simulated

Ideas

Go, reshape or park, with the evidence for and against.

  • › verdict: GO | RESHAPE | PARK
  • › red-team counter-case attached

App features

Which jobs a feature serves, where the gaps are, what it replaces.

  • › job served · gap · replaces
  • › against the same evidence base

Pricing and spend

Published prices, contract signals.

Competitors

Who's in the budget line you want.

Multi-model jury

No single model gets the last word.

Three models from three different labs play your buyers. A fourth argues against you. A fifth writes the verdict. A classifier re-checks the scores. No single model gets the last word, and if one is down, a named fallback steps in.

Auto-classification

Every finding, sorted before you read it.

Relevance, persona, job-to-be-done, stance and source, labeled automatically. The part that used to eat your evenings.

  • RELEVANT 0.88: Merchandiser · Mid-market · Pre-Book Gap Analysis · SUPPORTS · via Perplexity Agent
  • Single source (Podcast): brands still benchmark competitor prices by hand, over 3–4 weeks · via Firecrawl · Podcast
  • UNDERCUTS 0.89: published plans cap tracked competitors at 5–10 · via Perplexity Search
  • Vendor claim: labeled as marketing, weighted accordingly
  • UNSUPPORTED FIGURE: flagged, not in any source

Proof in numbers

Measured on real runs

7

passes per run

3

jury models

plus a red team, a judge and a cross-check

5–7

minutes per run

0

failed calls

104 calls in one run, 117 in the next

100%

cross-check

every jury score re-scored

3–5%

claims dropped

quote not found in the source

30

podcast results

in one run

41

podcast claims kept

in the same run

Old way vs ShipSift

The same decision, spent two ways

The old way

  • 20 browser tabs
  • Copy-pasting quotes into a spreadsheet
  • Asking one chatbot and hoping
  • Numbers you can't trace
  • A week to a decision

With ShipSift

  • One question
  • Cited, quote-checked evidence
  • A jury of competing models
  • Every number traced or flagged
  • A verdict in minutes
per run
5–7 min
each logged
7 passes
calls in the final ship check
0 failed
Talk to us

Trust

Built so you can check its work

Every number checked

Each figure in the memo must appear in a source behind the finding it cites, or it's flagged in plain sight.

No silent failures

If a step comes back short, the run says so and offers a one-click retry of only what's missing.

Evidence and simulation, kept apart

Cited findings are real; jury reactions are clearly labeled as simulated.

Fallbacks, named

If a model refuses or is down, a backup steps in and the run page says which.

FAQ

Questions people ask first

Tell us what you're trying to decide.

ShipSift is working with a small number of teams. Tell us the question and we'll take it from there.

Reply within 1 business dayNo spam, no mailing list.