
"Nobody asked for another tool. They asked to stop copying numbers between the ones they already had."
RemoteReps runs outbound sales teams for clients from a fully remote floor. In early 2026 that floor ran on Twilio for calls, a second dialer, Zapier for glue, Hubstaff for time tracking, HubSpot and Calendly for the agency's own sales, Meta Business for WhatsApp, and spreadsheets for everything the tools did not cover: daily rep sheets, client results, QA sheets. Managers computed decision-maker reach and appointment conversion by hand. Quality assurance meant a person listening to a handful of calls. The symptom was cost and lag. The real problem was that the most valuable work, listening to calls, writing training scripts, coaching reps, lived in people's heads, so no tool could hold it. Until models could read and listen, "one platform to rule them all" was a slogan. Now it was a design brief.
Seven tools, four kinds of user, and no shared truth. A single call existed in the dialer, in a manager's sheet, and in a client's emailed PDF, with three different statuses. Before designing screens, I needed to understand why the sprawl existed in the first place, because every tool had been added for a reason.
Every persona lived in a different tool. Reps lived in the dialer, managers in spreadsheets, clients in emailed PDFs, and executives in nothing at all. Nobody was looking at the same record, so every status meeting started with reconciliation instead of decisions.
KPIs were hand-computed, and therefore argued about. "Decision maker reached" meant something different on every campaign because each client counted a different call outcome as a win, and outcomes were typed as free text.
Integrations were the weakest surface. Zapier automations broke silently. A separate service on its own server relayed HubSpot and Calendly events into Slack and Google Sheets, and nobody owned it. The glue was the product people noticed most, for the wrong reasons.
The expensive work was unstructured language. Call reviews, training scripts, coaching notes, resume screening. None of it had a home in software because software could not do it, so it stayed manual, and manual meant sparse: a rep might get feedback on two calls a week.
Rather than four products, I designed one record of every call, lead, and appointment, then four views onto it: Executive, Manager, Client, and Rep. The hard part was not the layouts. It was deciding who is allowed to see what, and making that impossible to get wrong. Working with engineering, we settled on three scoping axes that every screen inherits:
The scoping lives in the data layer, not in a filter on the page, so a new screen cannot leak by accident. For single-campaign clients the campaign picker disappears entirely; the platform resolves it silently. Clients get a read-only portal with their own goal setting and scheduled reports instead of a PDF.
For each external tool we asked one question: does it hold data our users need to act on? If yes, it moved in. If it was a relay, it was retired. Each removal was framed as a UX decision, fewer logins, one notification surface, one activity feed, rather than an IT cost saving.
Call outcomes became a fixed vocabulary of twelve dispositions instead of free text. Each campaign declares which of them count as "decision maker reached" and names its primary outcome, whether that is a booked appointment, a qualified lead, or a sale. Every label on every dashboard adapts to that declaration, so the rep, the manager, and the client are reading the same number with the same name.
This is the decision that made the four dashboards one product. Once the number was shared, the arguments stopped, and goals could be set per day, week, and month against something everyone trusted.
"Consolidation is not a smaller toolbar. It is one definition of a good call that every screen agrees on."
Consolidation alone would have produced a tidy CRM. The work that made the agency valuable, coaching and quality, still lived outside software. This is the part of the project that could not have been designed two years earlier: the reason people used many tools was that the important work was human-only, and now it did not have to be.
Add a general assistant in the corner → it answers questions nobody was asking → novelty fades in a week and the actual workflow is untouched.
Let AI grade, message, and call with nobody accountable → one bad text to a lead or one unfair score → a client relationship or a rep's trust is gone.
Every recorded call is transcribed and scored on the agency's four pillars: opening, discovery, objection handling, and close. A manager writes a rubric as a sentence and the platform drafts the scoring criteria. Voicemails and robocalls are detected and left unscored, and an unusually high or low score triggers a second independent judge before a rep ever sees it. Coaching stopped being a weekly sample and became the default state of every call.
A new client's onboarding notes become a draft training script the account manager edits and the client approves through a public review link with inline comments. The approved script generates the rep quiz, and then the in-call navigator the rep follows while dialing. One source of truth flows from client intake to the rep's screen, where before it was three documents maintained by three people.
A live co-pilot reads the streaming transcript and surfaces the next objection handler before the rep needs it. An AI voice agent dials, qualifies, and warm-transfers to a human, and a manager can take over a live AI call from the same dialer. An AI SMS agent replies inside the unified inbox with a compliance check and an independent safety critic on every outbound message. Lead import maps any spreadsheet's columns automatically, and email drafts appear in the composer but never send themselves.
"The test for every AI feature was simple: does it produce the same object a person would, in the place the person already works?"
Every model call passes a cost budget gate, so a runaway feature degrades gracefully instead of surprising finance. Every AI voice conversation writes to an AI disclosure ledger, and the voice agent could only reach real leads after a staged certification from canary to general availability. Practice calls for coaching run through a separate browser channel so no telephony credential ever reaches a rep's machine.
And one rule held across all of it: the human owns Send. Drafts are drafts, suggested dispositions are suggestions, and a rep can always see why a score landed where it did.
A rep's quality score now comes from a model. That is a real risk to trust, and we did not pretend otherwise. We chose it because the alternative was a rep coached on two percent of their calls, which is not a fairer system, only a quieter one. The second judge on extreme scores, the unscored-call detection, and the visible reasoning on every score exist because we owned that tradeoff up front instead of discovering it in a complaint.
A platform this wide, forty-three modules and over a hundred screens in seven months, could not be built the traditional way by a small team. So the build process itself became a design problem, and AI was part of the answer there too.
Every feature runs through a pipeline of specialised agents, each with one job and one artifact: plan, research, UI design, implementation, adversarial review, fix, QA, UX audit, security, architecture, chaos testing, and a final ship or no-ship gate. Tasks are routed by risk, not size: a one-line change to data scoping gets the full independent review, while a copy fix does not. That rule was earned, not assumed. Twice, a tiny isolation fix that looked trivial carried a ship-blocking defect that only the review agent caught.
The same discipline applied to cost. When cloud CI for one month consumed three quarters of the organisation's budget, we moved the checks local and scoped them to what actually changed: from roughly a hundred runner-minutes per run to under three. Then, when the written rule alone did not change behaviour, we enforced it with a hook. The lesson from both: a rule nobody measures is a rule nobody follows, for agents as much as for people.
Call quality review went from a weekly sample to every call. KPI definitions went from per-manager spreadsheets to one shared vocabulary that reps, managers, and clients read the same way. The platform shipped in seven months across more than six thousand commits, with a live client running on it from the first quarter.
Consolidation is a UX decision, not an IT one. Nobody cares about a shorter vendor list. They care that the number on their screen matches the number on their manager's screen. Design the shared definition first and the tools fall away on their own.
AI belongs at the point of work, not in a corner. The features that stuck produce the same object a person would, in the place the person already is. The ones that would have failed were the ones that asked users to come to the AI.
Scoping is design material. Who sees what is the most important screen you will never draw. Putting it in the data model instead of the UI turned a hundred screens from a hundred risks into one decision.
Keep the human on Send. Every AI feature earned trust by stopping one step short of the consequential action. That constraint made the team comfortable shipping far more AI, not less.
Design the process, not only the product. The build pipeline, the risk-based routing, and the enforced rules were design work too. A product this wide is the output of a system, and the system deserves the same care as any screen.