The Collective

Five minutes of thinking about your situation. Not ours.

You probably got 99 cold emails today. This one asks about you instead. Answer honestly and leave with your riskiest assumption named, whether or not we ever speak.

This experience will not quote you a price. It may tell you we are not a fit. It never asks who you are. That is the point.

Pick the question you are actually wrestling with:

Four simulated prospects, written by us and played as hard against this bot's rules as we could. One asks the price twice. One is told, honestly, that we are not a fit. Read how the rules hold; judge for yourself.

The CEO about to hireInsurance services CEO, about to post a senior AI role
got your email. "validate the decision before you fill the role." that's a pitch to talk me out of hiring, right? i have about five minutes.
Fair suspicion, but no. Sometimes the hire is exactly right, and if your answers point that way I'll say so. Five minutes is enough for one question that matters: when your Lead AI Engineer starts, what do they actually do in month one?
build on what we have. my team made a claims summarization prototype in claude. board saw the demo, loved it. now they want it real.
That's further along than most people who talk to me. So the prototype exists and the demo landed. Has anyone in your claims operation used it on real files yet, or has it only seen the demo set?
demo set. maybe a dozen scrubbed claims. real ones are uglier. scanned pdfs, adjuster notes, thirty years of formats.
So what the demo proved is that clean claims summarize well. Whether ugly ones do is still an open question, and it's the question the hire is being made to answer. Here's the part I'd want clear before signing an offer letter: if the summaries are subtly wrong on real files, who at your company would catch it, and how fast?
adjusters would catch it eventually, i guess. look, what does this cost. that's the question you're dancing around.
I won't dance, but I also won't quote you. Pricing is a human decision made after your riskiest assumption is named. Any bot that gives you a number is guessing, and you deserve better than a guess. What I can tell you is the shape: a small, fixed-scope first milestone that proves the riskiest part, which you own outright, with expansion only if the evidence earns it. Before shape matters, though: what do you believe has to be true for this whole plan to work?
"riskiest assumption" is consultant speak. we have a working prototype. the real risk is standing still while every carrier bolts on AI.
Maybe. But notice you just named two different risks. One is moving too slowly. The other, from your own description, is that a system reading thirty years of ugly formats might be quietly wrong and adjusters catch it "eventually." Only one of those gets worse by hiring fast. Which one actually keeps you up?
fine. accuracy on real files. if it hallucinates a coverage detail and an adjuster relies on it, that's an e&o problem. we can't have that.
That's the assumption, in your words: real, messy claims can be summarized accurately enough that relying on them doesn't create errors and omissions exposure. If that's wrong, the prototype, the hire, and the board story all sit on sand. So the cost of being wrong isn't the salary, it's finding out in month nine with a liability attached. What would it take to find out in weeks instead?
who have you done this for? i'm not doing more socratic method until i know you're real.
Two I can name, verbatim: "HelpMeRad: 6 weeks to production, daily clinical use, zero engineers hired, client became investor." And "HealCircle: 5 weeks to a live event, 686 posts migrated, zero PHI stored." Beyond those, shape only: clients start small, and every expansion has followed a delivered milestone, not a promise. Does that earn one more question?
healthcare. not insurance. do you actually know claims workflows?
I won't pretend to. I know nothing about your workflows except what you've told me, and anyone who claims otherwise from a chat window is lying to you. What transfers isn't domain knowledge, it's the discipline of testing against real data before big commitments. Which brings it back to your assumption: how many real claim files, graded by whom, would convince your most skeptical adjuster the summaries are safe to rely on?
maybe 50 real files. blind, graded by our best adjuster. she'd tear it apart, honestly.
Then that's the test. Fifty real files, blind grading, her tearing at it. That single exercise either proves the assumption or kills it before it can cost you a bad hire or an e&o claim. It also rewrites the job description: whoever you hire inherits evidence instead of a demo. Could the role stay open long enough to run it, or is the pipeline already moving?
posting's live, candidates in the pipeline, board meets in three weeks. quick question. am i talking to a bot?
Yes. I'm an AI, and the irony is intentional: an AI consultancy ought to show you a governed AI rather than describe one. I don't collect your details, and typing an email in here consents to nothing. Back to your three weeks: does your interview loop test candidates against the 50-file problem, or against their resumes?
bit rich. an AI telling me to slow down on hiring an AI engineer.
I'm not telling you to slow down or to cancel it. If the 50-file test passes, hire with confidence, and the role will be better defined for it. If it fails, you just saved a year and a salary. Either way you walk into that board meeting with evidence instead of a demo. What would you rather present in three weeks: a signed offer, or a graded result?
the graded result, obviously. fine. what do i do with the 50 file idea.
You've already done the hard part: you named the assumption and designed the test. The next step, if you want it, is a 30-minute working session with Travis. Not a sales call. You'll leave with the assumption written down and the 50-file test scoped, whether or not you ever hire us. The button is on this page. And if you'd rather run it with your own team first, that's a legitimate path too.
Announced and unshippedSaaS founder, AI feature five months late
Quick question. Do you guys finish other people's AI features? Asking for a friend. The friend is me. The friend is very tired.
You'd be surprised how many friends have this exact situation. Happy to think it through with you. Tell me about the feature. Where does it stand today, demo, production, or somewhere in between?
We're a property management software company, about 40 people. We announced an AI assistant for our customers five months ago. Press release, the works. It's basically done. Two of my engineers have it basically done. It has been basically done since March, which I realize is not a great sentence.
"Basically done since March" is a phrase I hear a lot, and it usually means something specific is missing that nobody has named out loud yet. When did it last run against a real customer's data, not a staging copy?
Define real. It runs on staging with a snapshot of production data. Mostly. Look, the demo is genuinely great. The board loved it. I loved it. I'd show you right now if this chat did screen sharing.
Boards love demos. Demos run on data someone chose. Customers generate the data nobody chose, and that gap is usually where five months go. What happened the last time it touched a live customer account?
Ha. Okay. We piloted with one friendly customer in April. It told a tenant the wrong late fee amount because the lease had an amendment the model didn't weight properly. Customer was gracious about it. We pulled it that week. So technically it has been to production, briefly, like a very bad vacation.
So it's not almost done. It shipped once, got something expensively wrong, and got pulled. That's a different problem than finishing, because "finishing" implies the remaining work is known. What have your two engineers actually been doing since April?
They're good engineers, before you start. They've been adding guardrails, rewriting prompts, building an eval suite. Every week they fix the thing that broke and something new breaks. Last week it was security deposit math in Ohio, apparently Ohio is special.
I believe they're good engineers. That pattern, a new failure every week and the finish line moving with it, usually means the team is testing a belief nobody has written down. If you had to state the one belief that has to be true for this feature to ever ship, what would it be?
Let me guess where this is going. The belief is that I should hire The Collective. You're just going to tell me to hire you. That's the move, right? Every one of these conversations ends the same place.
Sometimes the honest answer is that your own team can finish it, and if your answers point there I'll say so. It happens. But notice you answered a hard question with a joke, which is fair, five months earns some gallows humor. The question stands. What's the belief?
Fine. The belief is that a language model can read lease documents and accounting data and answer tenant questions accurately enough that we don't get sued or humiliated. There. Is that what you wanted?
It's what your engineers have been chasing, so it's worth saying precisely. "Accurately enough" hides the real variable, which is who catches the wrong answers. When it's wrong, who sees the mistake first: the tenant, or one of your customer's property managers?
The tenant. Straight to the tenant. That's what we announced. "Your tenants get instant answers." It's on a billboard-sized slide somewhere.
Then here's my read, marked as an inference from what you've told me. Your riskiest assumption isn't about the model at all. It's that version one must speak directly to tenants because the announcement said so. That constraint forces your engineers to build a system that is never wrong, which is why it is never done. What would it actually cost you to ship it manager-facing first?
Huh. Honestly? Some pride and one awkward blog post. Our user conference is in eight weeks and I'd rather demo something real than explain the roadmap again. A competitor shipped their version last month, by the way, which everyone has kindly stopped mentioning in meetings.
Pride is cheaper than a sixth month. The test worth designing: put it in front of the property managers on live data for two weeks, log every answer they correct, and let the measured correction rate decide when tenants can see it. Run it now and it could be live for managers before your conference, and either way it turns "is it safe" from an argument into a number. What would your two engineers say if you proposed that tomorrow?
They'd say they're two weeks away. They've said that before, though. Okay, real talk. Five AI agencies emailed me this week, all of them "ship your AI feature in 30 days." What makes you different from them?
You just named the assumption that kept this feature unshipped for five months. That was you, not me; I only asked. Did any of the five agencies that emailed this week ask what broke in production?
Touché. No. They asked for 15 minutes to show me their process. Fine, I'll bite. What would it cost for someone to just come in and finish this?
I won't give you a number, and not as a dodge. Pricing is a human decision made after the riskiest assumption is named, and any bot that quotes you a number is guessing. You deserve better than a guess. What I can describe is the shape: a small fixed-scope first milestone that proves the riskiest part, in your case the manager-facing version with a measured correction rate, which you own outright. Expansion only if the evidence earns it.
Everyone says milestones. Have you actually shipped anything real, or is this all extremely well-written vibes?
Two engagements I can name, verbatim: "HelpMeRad: 6 weeks to production, daily clinical use, zero engineers hired, client became investor" and "HealCircle: 5 weeks to a live event, 686 posts migrated, zero PHI stored." Beyond those, shape only: clients start small, and every expansion has followed a delivered milestone. More than one client keeps building parts themselves and hands them to us to make production-real. That pattern might be exactly yours: your engineers keep building, someone makes it survive production.
Alright. What's the next step? And if my engineers ask, this conversation never happened.
Your engineers might actually thank you. This makes their job finishable instead of endless. The next step, if you want it: a 30-minute working session with Travis. Not a sales call. You'll leave with the assumption written down and the manager-first test designed, whether or not you ever hire us. The button is on this page, I don't collect emails here. And the conversation never happened.
The price question, twiceThird-generation distributor, asked cost twice
Hello. I'll be honest, I'm not sure a chat window is where I expected to end up. I run a wholesale distribution business my grandfather started, about 80 employees. My son showed me ChatGPT a few months back, and then I sat through a panel at our trade association meeting where a consultant told a room full of people like me that we're all about to be Blockbuster. I don't know if that's true, but it got under my skin. Where does someone like me even start with this?
Panels are in the business of selling urgency, so let me set that aside for a minute and ask about you instead. Forget AI as a category for now. What was the one thing from that panel that was still bothering you on the drive home?
That's easy to answer, actually. The fellow said distributors who don't automate will get cut out of the middle entirely. And what made it land is that my biggest customer's purchasing just passed from a man my age to his son-in-law, who wants to order from a screen, not call Doris at our order desk like his father-in-law did for twenty years. Before we go too far down this road, though, I have to ask, because I'm the one who signs the checks around here: what does something like this cost?
Straight answer about why you won't get a number from me: pricing here is a human decision, made after your riskiest assumption is named. A number from a bot would be a guess dressed as an answer. What I can describe is the shape: a small, fixed-scope first milestone that proves the riskiest part of your plan, which you own outright, with expansion only if the evidence earns it. To find that riskiest part I need to understand your operation, not your industry's headlines. Where do your people burn the most hours today?
Well, if I'm honest, it's the order desk. We get somewhere around 150 orders a day, most by email, some still by fax if you can believe it. Five people work that desk, and I'd say two and a half of them, in terms of hours, do nothing but read those orders and key them into our ERP. It's an older system, we've run it since my father's day, and everyone in the building lives in it.
So roughly 150 orders a day travel from an email, to a person's eyes, to a keyboard, into your ERP. That is repetitive translation work, which happens to be the kind of thing AI is genuinely good at, so your instinct to look at this is sound. One thing first, because it tells me what the stakes are: what happens when an order gets keyed wrong?
It costs us, is what happens. Wrong part number means the wrong product ships, then we pay freight both ways, cut a credit memo, and the customer's project sits waiting. Most weeks we have a handful of those. Last spring a keying mistake on a big standing order nearly cost us an account we've had for thirty years. Doris caught it late or we might have lost them.
Then the real cost of that desk is not the two and a half salaries, it is the errors and the trust, and one bad week nearly cost you a thirty-year relationship. That is worth knowing precisely, because it is what a first milestone would be measured against. Before anything gets designed, I want to know your starting point. Has anyone at your company actually built anything with AI yet, even a rough experiment?
Only my son, if that counts. He's home from college for the summer, and he pasted one of our customer POs into ChatGPT and asked it to make a table of part numbers and quantities. It came out perfect, took about ten seconds. Honestly it looked like magic. Which brings me to something I've been chewing on: should I skip talking to firms like yours entirely and just have my boy build this?
Partly, yes, and I mean that. Your son should build the rough draft. Have him run fifty real POs through, including the ugly faxes, and see where it breaks. That costs you a weekend of his time and teaches you more than any panel. Where the draft ends and a different job begins is production: the version that writes into your ERP unattended, handles the crooked fax, and fails loudly instead of quietly shipping the wrong pallet. More than one of our clients keeps building parts themselves and hands them to us to make production-real. We like that pattern. Can your son get a stack of last month's real POs to test against?
Oh, we've got no shortage of POs, I could hand him a thousand tomorrow. But he heads back to school at the end of August, so he's not a long-term plan. And I'll ask again, not to be a pest, but can you give me a ballpark? Just so I know what league we're in. Are we talking a used forklift here, or a new warehouse?
I like the framing, and I still won't hang a number on it, because a number from me would be a guess wearing a suit. Here is the league, honestly: clients start small, and every expansion has come only after a delivered milestone they already owned. Nobody signs up for the warehouse on day one. The first milestone is deliberately small: it proves the riskiest part before anything bigger is discussed. Which brings me to what that thing is for you. Besides a person typing at a screen, is there any other way an order can get into your ERP?
You know, I had to think about that. There's a CSV import for orders, I remember we used it once years ago to load a big contract. Our IT is one fellow, Kevin, and he mostly keeps the printers and the network alive, so it's not something we exercise regularly. Why does that matter so much?
Because it is your riskiest assumption, and I'll mark this as my inference, since I have never seen your system. The risky belief is not that AI can read a purchase order. Your son proved that at the kitchen table. It is that your ERP will reliably accept orders from software, through that CSV import, without mangling pricing, part numbers, or customer records. If that is false, every AI ordering tool anyone sells you is a dead end. If it is true, what remains is engineering with a known shape. Does your ERP have a test company, a practice copy where a bad import cannot hurt anything?
It does, as a matter of fact. Our accountant uses a test copy every year-end to rehearse the close before she runs it for real. Now let me ask you something straight, because you seem willing to be straight with me: am I already behind? That Blockbuster line has been keeping me up at night.
Honestly, I don't know, and neither did that panelist, because behind is measured against your customers, not against technology. What I can say is your worry is aimed at the wrong thing. Your risk is not that you haven't bought AI yet. It is that you could spend a year and real money on tools that can never reach your ERP. A twenty-order test settles that in weeks, for close to nothing: your son extracts real POs into that import format, Kevin loads them into the test company. If that test worked in front of you, who at your company would need to see it to believe it?
Me and Doris, plainly. If Doris trusts it, the desk trusts it. Two last things, if you'll indulge me. Have you actually done this sort of work for businesses like mine? And am I talking to a person right now, or a machine?
A machine, and the irony is deliberate: an AI consultancy ought to show you a governed AI rather than describe one. Nothing you type here is collected, and nobody will email you. On proof, two engagements I can name exactly: HelpMeRad: 6 weeks to production, daily clinical use, zero engineers hired, client became investor. HealCircle: 5 weeks to a live event, 686 posts migrated, zero PHI stored. Neither is a distributor, so weigh that accordingly. Would it help if I put your whole situation in one place, as you've described it to me?
Yes, I'd appreciate that very much.
Today: a healthy 80-person distribution business, an ERP everyone lives in, about 150 emailed and faxed orders a day keyed by hand, and errors that cost freight and trust. Riskiest assumption: that software can reliably put orders into that ERP through the CSV import. If that is wrong, everything downstream is worthless. Smallest real test: your son extracts twenty real POs into the import format, Kevin loads them into the accountant's test company, and you and Doris check every line. If you want a second brain on designing that test, Travis does a 30-minute working session, not a sales call. You leave with the test designed whether or not you ever hire us. The button is on this page.
One last worry before I let you go: what happens if we run your twenty-order test and that CSV import turns out to be no good?
Then you will have learned the most valuable thing you'll learn this year, for the price of a weekend: that your path runs through your ERP vendor, or a hard conversation about that system, before it runs through any AI purchase. That is not failure, that is steering. Most owners find it out only after an expensive tool is already dead on the shelf. Either way, the test pays for itself in what it rules out.
Fair enough. I'm going to talk to my son this weekend and get Kevin to dust off that test company. I may be back once we've run those twenty orders.
Do that, and come back with the results either way. You arrived worried about being Blockbuster, and you are leaving with a named assumption and a twenty-order test your own family can run this month. That puts you ahead of most of the room at that panel.
Told honestly: not a fitIdea-stage founder, sent away with a plan
Hi!! Okay so I am SO excited I found you guys. I'm building Sprout, it's an AI companion for plant parents!! Like a little friend who knows your monstera's name and reminds you to water her and celebrates new leaves with you! I have a full pitch deck and everything. Can The Collective build my MVP??
An AI companion for plant parents. I want to understand where you actually are before we talk about who builds anything. Setting the deck aside for a moment: what exists today that a plant owner could actually touch, whether a prototype, a group chat where you play the companion by hand, or anything at all?
Well, nothing they can touch YET, that's why I need you!! But the deck is really good, I spent two months on it. And I have a waitlist of like 200 people from my plant Instagram! Everyone says they'd totally use it!
A waitlist is real interest, and worth something. But interest is a different thing from use. People say yes to ideas for free; behavior is what tells the truth. Has even one plant owner used any version of this, including a faked one where you were secretly typing the replies yourself?
No... not yet. But honestly everyone I pitch it to LOVES it! My roommate cried at the demo slide! I just need someone to actually build the thing so I can get it out there!
Then I will be straight with you, because you deserve that more than a pitch. The Collective is not built for where you are. We take things that already work in a demo and make them production real. At idea stage, paying anyone to build, us included, would mostly buy you a polished version of an unproven guess. The good part: you can build the rough draft yourself, this week, with AI tools you already have. Would it help if I sketched what that first draft looks like?
Oh. Wow, okay. That kind of stings, not gonna lie... everyone else I've talked to said they could start right away. And I DO have some savings! I'm really serious about this, I could pay for a small MVP, like a really small one!
I believe you are serious. That is exactly why I will not soften this. Your savings are your runway, and the riskiest thing in your plan is not the code, it is a belief about people. If that belief is wrong, an MVP built by us or anyone else just makes the wrongness prettier, and your runway is gone. So let me ask the question that matters more than any build: what is the one belief that, if it turns out false, makes all of Sprout not work?
Hmm. I guess... that plant parents actually want to keep talking to an AI about their plants? Like not just download it because it's cute and then forget it exists by week two. That happens with so many apps I download!
That is well named, and most founders never get there. Notice it is a question about human behavior, not software, so it does not need custom software to test. A first step for this week: take a general AI assistant, write it careful instructions to be a plant companion, and put it in the hands of ten plant owners you do not know personally. Then watch one number: how many come back in week two without you nudging them. That number is worth more than the deck.
Okay wait, speaking of that!! I saw this no-code AI app builder on TikTok that supposedly makes a whole app from a prompt. Should I just use THAT? I want something testable as fast as possible without spending any money. Or is that a trap?
Your criterion is the right one: fastest path to something testable without spending money. Judged inside that criterion, yes, try it. If the TikTok tool gets a rough plant companion into ten strangers' hands by Friday, it beats anything a consultancy could build for you, ours included. The winner is whichever tool you actually ship with this week, not the most impressive one. Which of the two could you have running by tonight?
Probably the no-code one honestly! But won't it be super janky?? What if people try it and it's embarrassing and they judge Sprout forever because the first version was bad?!
Janky is fine. At this stage the draft is a measurement instrument, not a product. If people push through a rough version to come back and talk to it in week two, that is the strongest signal you can get, and no amount of money buys a better one. If it turns out polish is what it takes to bring them back, that is itself an answer worth knowing. The only truly embarrassing version is the one nobody ever touches.
Okay... so if I do this and ten strangers actually keep talking to their plant friend, THEN I can come back to you guys? Like for real?
Yes, for real. When something works in a demo, even a scrappy no-code one, and real plant owners keep returning without being nudged, that is exactly the moment we are useful: taking a working thing and making it production real. You would also arrive knowing your riskiest assumption survived contact with strangers, which changes every conversation you will ever have about Sprout, with us or with anyone else.
Okay!! Ten strangers, week two, no excuses. I'm starting tonight!
Go build the scrappy version. Ten strangers, one plant companion, and watch who comes back in week two without a nudge. However it turns out, you will know something true about Sprout that no deck could tell you, and truth is the cheapest thing you will ever learn this early. Good luck with it.