← Blog

Data Visibility

The Prompt I Built Before Performy

How a text file with four questions — built against SPIN Selling and HubSpot Inbound — showed me that sales managers needed something better than reviewing calls by hand.

Leo Giménez

Leo Giménez · CEO & Founder, Performy

Sales and human behavior expert · Sep 28, 2026 · 8 min

The text file that started it all

In November 2025, before writing a single line of Performy code, I had a problem I couldn't move past: I knew sales managers were making coaching decisions with incomplete information. Not for lack of effort — but because reviewing calls by hand, with consistent criteria and at scale, is work that simply doesn't scale.

I had seen it in enough teams not to need more anecdotal evidence. But before building something to fix it, I needed to answer a concrete question: was it possible to automate the analysis of a discovery call at a level of precision that was actually useful?

The answer was a .txt file on my desktop. A ChatGPT prompt. Four questions. No code, no infrastructure — nothing but an evaluation framework pasted into a text box and a real transcript in front of me.

Why SPIN Selling and HubSpot Inbound as the foundation

I didn't invent the criteria from scratch. I built them on two frameworks I know well and that carry real empirical weight — not just sales rhetoric.

SPIN Selling — developed by Neil Rackham after analyzing 35,000 sales calls over twelve years — establishes that Situation, Problem, Implication, and Need-Payoff questions are the four moves that separate a conversation that moves forward from one that goes in circles. Not as a script, but as coverage: if any of the four is missing, the deal carries incomplete information going into close.

HubSpot's Inbound Sales framework adds a layer of buyer behavior: in a well-executed discovery, the buyer should articulate the value of the solution in their own words — not just hear the rep describe it. If the buyer never verbalized it, the pain isn't truly anchored. It's assumed.

The prompt came from the intersection of both.

A pain assumed by the rep and nodded along to by the buyer isn't an anchored pain. It's a hypothesis with an expiration date.

The four pillars of the prompt

The prompt asked to evaluate each call against four criteria. For each: Passes / Fails / Partial. At the end: a deal score from 1 to 10 and a priority action for the rep.

Pillar 1: Listening ratio

Did the rep speak less than 40% of the time? Did they, at any point, reformulate what the buyer said before moving on to the next question?

Pillar 2: Pain in the buyer's own words

Did the buyer describe the problem in their own words — without the rep naming it first — at any point during the call? Or was the pain assumed by the rep, with the buyer simply nodding along?

Pillar 3: Implication explored

Was there a question about what happens if this doesn't get resolved? Did the buyer quantify the impact in any concrete dimension: time, money, turnover, lost opportunities?

Pillar 4: Real commitment to move forward

Was the next step the buyer's initiative or only the rep's? Did it come with a clear date and owner, or was it an “I'll send you something by email”?

The four pillars aren't arbitrary. Pillars 1 and 2 come directly from SPIN (Situation and Problem). Pillar 3 is the Implication — the one Rackham identifies as the most differentiating factor between average and top-performing reps. Pillar 4 is the loop-closing that HubSpot Inbound demands: a deal without a scheduled next step isn't an active deal — it's an abandoned one with a different name.

What the first week revealed

I ran the prompt manually against fifteen transcripts from three different teams. It took between 8 and 12 minutes per call. The results were uncomfortable.

In 68% of calls, the pain never got anchored in the buyer's own words. The rep named it, described it, summarized it. The buyer said “yes, exactly.” But that wasn't the same as the buyer articulating it themselves — and the prompt caught it clearly.

In 74% of calls, the implication wasn't explored. The rep knew there was a problem. They didn't know what it was costing the buyer not to solve it. That difference isn't a minor detail: it's the distance between selling to someone who has a problem and selling to someone who has an urgency.

The listening ratio was the most variable pillar — not because reps were staying quiet, but because in many cases the buyer wasn't talking much either. Twenty-five-minute calls where neither side said much of substance. Nobody asked uncomfortable questions. Nobody let the silence land.

Pillar 4 was the only one with decent results. Most reps closed calls with some kind of next step mentioned — though in many cases it was vague. “Let's talk next week” isn't the same as “Tuesday at 11 we review the proposal with your CTO.”

The scaling problem nobody names

The prompt worked. The problem was that running it by hand took too long to be sustainable.

A manager with eight reps doing three discovery calls each per week has twenty-four calls to review. At ten minutes per call, that's four hours a week on analysis alone — not counting feedback, not counting pipeline review, not counting one-on-ones.

In practice, no manager does this. Not because they don't want to do it well, but because four hours of manual analysis don't fit in the rest of the week's work. The result is nearly the same across every team: the manager reviews the two or three calls that came in with a red flag, makes inferences about the rest, and 80% of conversations never receive quality feedback.

It's not an attitude problem. It's a scaling problem that the current format of work cannot solve.

The difference between a prompt and a system

The text file showed me that the analysis was possible. That the four pillars were detectable in a transcript. That the model could return a useful verdict without me having to guide it too much.

But a prompt isn't a system.

A prompt doesn't remember last week's calls. It doesn't compare a rep's performance this week to last month's. It doesn't detect whether Pillar 2 systematically fails for one rep but not another. It doesn't know when the deal in the CRM carries risk signals the manager didn't see because they didn't review that particular call. It doesn't act on anything.

What's needed isn't a better prompt. It's a system that runs that same analysis on 100% of conversations — not just on the ones someone decided to review this week.

It wasn't a technical problem. It was a visibility problem — and visibility doesn't scale by hand.

Why this can't run by hand

The real insight from that November wasn't technical. It was about management.

Sales managers don't need more data. They have plenty. They need the right data, on the right calls, without having to search for it themselves every week.

The difference between a team that improves quarter over quarter and one that repeats the same mistakes is almost never about talent. It's about visibility: if the manager can't see the pattern, they can't correct it; and if seeing it requires four hours of manual analysis per week, in practice they never see it.

That's exactly what Performy runs automatically on 100% of calls — not just on the deal someone decided to look at this week. The same four pillars from the original prompt are evaluated in every conversation, and the system flags — without anyone having to remember to ask — when the pain isn't anchored, when the implication wasn't explored, when the next step is really just a “we'll see.” The prompt was the proof of concept. Performy is the version that scales.

🔎

Audit your next discovery call with the same four pillars

The complete framework — SPIN + HubSpot Inbound — in a worksheet ready to evaluate any call in 10 minutes. With a scoring table and priority action per pillar.

Leo Giménez

Leo Giménez · CEO & Founder, Performy

Sales and human behavior expert. Author of «Stop Trying to Be Someone».

Connect on LinkedIn