Formerly IdeaMan · 7-pass multi-model research engine
Research that argues with itself, so your roadmap doesn't have to.
ShipSift runs your question through a 7-pass, multi-model research pipeline. It gathers cited evidence from the web and podcasts, labels every finding by persona and stance, puts it in front of a jury of competing AI models, and hands you a verdict you can defend.
Models and sources ShipSift orchestrates
One question. Models and sources from:
Trademarks belong to their owners. ShipSift is not affiliated with or endorsed by them.
How it fits together
One question in, one decision out
Ask about a company, a persona, an idea or a feature. ShipSift fans out across deep web research, targeted search, full-page reads and podcast transcripts, then converges on one memo.
The 7-pass workflow
Not a chatbot: a pipeline of separate stages, each logged with its own health report.
01Research
01Research
A deep web research agent finds what's actually been published: prices, customers, case studies, job postings. Figures missing from snippets get re-checked on the full page.
- perplexity-agent
02Evidence
02Evidence
Targeted searches, page reads and podcast episodes, with a quote pulled for every claim. A claim whose quote isn't in the source is dropped.
- search
- page reads
- podcasts
03Triage
03Triage
A classifier labels each finding by relevance, persona, job and stance, and flags any figure its source doesn't contain.
- relevance
- persona
- job
- stance
04Jury
04Jury
Three different AI models react in character as each of your buyer personas.
- 3 models
- 3 labs
05Red team
05Red team
A separate model builds the strongest case against the idea.
- counter-case
06Judge
06Judge
A frontier model writes the memo: a verdict per persona, a scorecard, and the questions to ask real people.
- go / reshape / park
07Cross-check
07Cross-check
An independent classifier re-scores the jury. Disagreements are shown, not hidden.
- independent
Live run log
Every run leaves a receipt
This is a replay of a real run: what ran, what it kept, what it dropped and why.
$ shipsift run --persona "Merchandiser · Mid-market" --question "Would $5M+ Shopify apparel brands pay for sourced competitor intel?"
▸ research perplexity-agent · 3 pages re-checked ✓ 1m 37s
▸ evidence 5 queries → 5 results · 1 page read ✓
▸ alexandria 3 queries → 30 podcast results
asked: "competitor price monitoring software for apparel brands" …
▸ podcasts 30 pieces → 44 claims · 41 kept · 3 dropped (quote not found)
▸ extraction 57 claims · 3 dropped (5%) · 54 kept ✓ 1m 1s
▸ triage relevance · persona · job · stance · 1 figure flagged ✓
▸ jury mistral · gpt · grok × persona 3/3 ✓
▸ red team counter-case written ✓
▸ judge memo · verdict: RESHAPE · number check clean ✓
▸ cross-check 15/15 re-scored ✓
run complete · 104 calls · 0 failed · 5m 31sWhat you can research
One evidence base, four kinds of question
Research companies, buyer personas, product ideas and app features side by side, against the same evidence.
Companies
Who they sell to, what they charge, who uses them, sourced.
- › customers · pricing · case studies
- › every line carries its quote
Personas
How merchandisers, marketers, designers or any role actually work, and what they'd pay for.
- › job-to-be-done · budget line
- › reactions labeled simulated
Ideas
Go, reshape or park, with the evidence for and against.
- › verdict: GO | RESHAPE | PARK
- › red-team counter-case attached
App features
Which jobs a feature serves, where the gaps are, what it replaces.
- › job served · gap · replaces
- › against the same evidence base
Pricing and spend
Published prices, contract signals.
Competitors
Who's in the budget line you want.
Multi-model jury
No single model gets the last word.
Three models from three different labs play your buyers. A fourth argues against you. A fifth writes the verdict. A classifier re-checks the scores. No single model gets the last word, and if one is down, a named fallback steps in.
Auto-classification
Every finding, sorted before you read it.
Relevance, persona, job-to-be-done, stance and source, labeled automatically. The part that used to eat your evenings.
- RELEVANT 0.88: Merchandiser · Mid-market · Pre-Book Gap Analysis · SUPPORTS · via Perplexity Agent
- Single source (Podcast): brands still benchmark competitor prices by hand, over 3–4 weeks · via Firecrawl · Podcast
- UNDERCUTS 0.89: published plans cap tracked competitors at 5–10 · via Perplexity Search
- Vendor claim: labeled as marketing, weighted accordingly
- UNSUPPORTED FIGURE: flagged, not in any source
Proof in numbers
Measured on real runs
7
passes per run
3
jury models
plus a red team, a judge and a cross-check
5–7
minutes per run
0
failed calls
104 calls in one run, 117 in the next
100%
cross-check
every jury score re-scored
3–5%
claims dropped
quote not found in the source
30
podcast results
in one run
41
podcast claims kept
in the same run
Old way vs ShipSift
The same decision, spent two ways
The old way
- 20 browser tabs
- Copy-pasting quotes into a spreadsheet
- Asking one chatbot and hoping
- Numbers you can't trace
- A week to a decision
With ShipSift
- One question
- Cited, quote-checked evidence
- A jury of competing models
- Every number traced or flagged
- A verdict in minutes
- per run
- 5–7 min
- each logged
- 7 passes
- calls in the final ship check
- 0 failed
Trust
Built so you can check its work
Every number checked
Each figure in the memo must appear in a source behind the finding it cites, or it's flagged in plain sight.
No silent failures
If a step comes back short, the run says so and offers a one-click retry of only what's missing.
Evidence and simulation, kept apart
Cited findings are real; jury reactions are clearly labeled as simulated.
Fallbacks, named
If a model refuses or is down, a backup steps in and the run page says which.
FAQ
Questions people ask first
Tell us what you're trying to decide.
ShipSift is working with a small number of teams. Tell us the question and we'll take it from there.