AI that unlocks your team's highest and best use.

Every hour a senior person spends retyping what a PDF already says is judgment you paid for and didn't get. We build the systems that take that work off your team — inside your operation, in ninety days — so the same payroll carries more volume, and each additional job, claim, or account costs less to serve than the last. That is operating leverage: the direct result of removing work that never needed a person.

01Where it breaks

AI doesn't fail on the model. It fails on everything around it.

01

Nobody defined the problem.

The tool got chosen before anyone asked what the business was actually trying to do. So it automates a workflow nobody questioned, inside a process that exists because of a decision made in 2011. Faster is not better. The fastest way to do the wrong thing is to automate it.

02

Complexity got sold as sophistication.

There is money in making this look harder than it is. An agent framework where an off-the-shelf tool would do. A model where an incentive change would do. A data warehouse for a company with eleven thousand rows. Ask what it would take to remove a piece — if nobody can answer, it was built to impress, not to work.

03

The team was never brought along.

Something got installed and a training deck got emailed. Six months later people are back on the spreadsheet, because nobody sat with them, nobody adapted it to how they actually work, and nobody was ever named as the owner. Adoption isn't the last phase of the project. It's most of the job.

03Standing rules

These aren't values. They're operating constraints.

  1. 01Business problem first. Technology second.
  2. 02We promise exactly what we can do. Nothing beyond it.
  3. 03If something proven already solves it, we'll name it and step back.
  4. 04No jargon as a moat. Everyone on your team should be able to question what we build.
  5. 05We build around how your team already works — or we change how they work deliberately, having said so out loud first.
  6. 06It ships into production, or we don't build it. No pilots.
  7. 07You own the system, the code, the credentials, and the roadmap. From day one.
  8. 08Nothing gets ripped out. We integrate at the edges.
04Why now

The head start is the advantage.

One workflow. Boring, high-volume, reversible.

Then the function around it — same data, same habits.

Each one costs less than the last.

01

The learning is the asset.

The first system teaches you which work actually needed a person, where your data was wrong, and what a process really costs to run. That knowledge is worth more than the software, and it isn't purchasable — you can only start earning it. A competitor who begins next year begins at zero on all of it.

02

Operating leverage, not cost savings.

Hours returned to a good team rarely show up as savings — they show up as more quotes out, more jobs closed, more accounts served on the same payroll. That is the definition of operating leverage, and in a service business it is the most durable margin advantage available: your cost base stops moving in lockstep with your revenue, and every additional job is worth more than the one before it.

03

The second one is much faster.

Once the data is clean, the systems are connected, and your team is fluent, the next department takes a fraction of the time and cost. The firms pulling ahead in your industry aren't ahead because they bought better technology. They're ahead because they already paid the setup cost, and they're now spending it once instead of every time.

Limited engagements.

We take a small number of engagements at a time, because the work is hands-on and done inside your operation. Tell us what's breaking. If the answer is a product you should just go buy, we'll tell you that on the first call and point you at it.

Request an engagement
01The method

The order is the point.

Why fixed order

Every engagement runs the same four phases in the same order. The discipline is the product — it's why the systems are still running at month twelve while the pilot graveyard fills.

Phase 01

Survey

2 weeks

We start with the owner, not the org chart. What are you actually trying to do — this year, and in five? Grow without adding headcount? Get out of the day-to-day? See the numbers more than once a month? Sell in three years to someone who'll pay for a business that runs without you? Every technical decision downstream depends on that answer, and it's the question most implementation firms skip.

Then we learn the work. How many touches a job takes, who's involved at each one, where it waits, where it comes back for correction, what gets typed twice. We sit with the people doing it — because the person who's done it for nineteen years knows things about your business that have never been written down, and often can't articulate them until someone asks the right question.

Output: a written map of the operation and a ranked roadmap of what's worth doing, with the reasoning attached. That document is yours. You can hand it to us, hand it to someone cheaper, or run it yourself. We'd rather you have a real artifact than a dependency.

Phase 02

Decide

1 week

Four options for every candidate on the list, and they're genuinely weighted equally:

  • Buy it. Something proven already does this. We'll name it and tell you what it costs.
  • Build it. Nothing off the shelf fits how you work, and the volume justifies a custom system.
  • Change the process. It's an incentive, a handoff, or a policy — and no software will fix it.
  • Leave it alone. The volume is too low, the process is about to change, or it just isn't worth it.

The build-versus-buy call turns on a question nobody asks out loud: how much behavior change can your team absorb right now? A proven tool your people won't use is worse than a rough one shaped around what they already do. If the crew won't file mobile expense reports but will happily text a photo of a receipt, that's not a discipline problem to be solved with a policy — it's a design constraint, and there's a real decision to make about which side gives.

We also calibrate to how much risk you want. There's a frontier of what's technically possible and a nearer edge of what's commercially sensible for your industry and your tolerance. We'll tell you where both are, and build to the one you choose.

Phase 03

Deploy

6–8 weeks

One thin path, end to end, into production. Not a demo environment.

A human reviews every output until measured accuracy earns the review down to sampling. We instrument before we trust: every run logged, every correction captured, error rate on a dashboard you own.

Existing systems stay up. We integrate at the edges. Nothing gets ripped out, and nobody is asked to learn a new way of working unless we've agreed together that the new way is worth the disruption.

Phase 04

Hold

90 days, then a decision

We stay through the first 90 days of real load, because that's when the exceptions arrive — drift, edge cases, seasonal volume, the statement format that changes in January.

This is also when the teaching happens. Not a training deck: sitting with the people who use it, in their work, until they can run it, extend it, and explain it to the next person. Where an owner or an executive wants to work this way themselves, we coach that directly — what to use, when not to, and how to tell the difference.

We write the runbook and name an internal owner. At 90 days you take it in-house or retain us. Both are real options, and we'll tell you which we think you should pick.

The part most firms won't say

What we'll tell you not to do.

A

Something already does this.

If a proven, well-supported product solves your problem, we'll name it, tell you what it costs, and help you set it up properly. Custom software you have to maintain forever is not a prize.

B

It's an incentive problem.

If the reports are late because nobody is measured on them, no system will fix that. We'll say so, and the conversation becomes a management one. Some of the highest-return changes we've recommended cost nothing to implement.

C

The process isn't stable.

If how you work changes every quarter for real business reasons, automating it now destroys money. Stabilize first, then automate. We'll help you decide which.

D

Nobody will own it.

If no one internally will own the system after handover, don't start. It decays in six months, and you'll conclude the technology doesn't work — which is the wrong lesson, and an expensive one to learn.

The survey is where it starts.

Two weeks inside your operation. A ranked roadmap of what's worth doing — and what isn't. Scope comes after, not before.

Request an engagement
02Capabilities

Three things happen at once.

No packages

Transformation isn't a product you buy once — it's assessment, building, and teaching running together. These aren't tiers and there's no ladder to climb. The mix depends on what your business actually needs, which is what the survey is for.

Mode 01

Assess

Understanding the business well enough to know what's worth doing. Owner objectives, current workflows, touch counts, gaps, failure points, and where the money and the time actually go.

Produces a ranked roadmap with the reasoning attached — a real document you own and can take anywhere, including to someone else.

Mode 02

Build

Systems, agents, integrations, and automations — plus the plain digital work that often matters more.

Document & data extraction
Turning statements, invoices, work orders, permits, quotes, and field forms into structured records, with validation rules, confidence thresholds, and review queues for anything uncertain.
Workflow integration
Connecting systems that were never designed to talk. CRM, billing, email, dispatch, spreadsheets — including the reconciliation logic and exception handling that decide whether it still works in month six.
Reporting & visibility
Most owners find out how the business did when the books close, weeks after they could have done something about it. Getting to a weekly picture is mostly a data-collection and incentive problem, and it's frequently the highest-return thing on the list.
Internal retrieval
Answering questions from your own documents: policies, historical quotes, SOPs, contracts. Cited to source, or it doesn't answer.
Scan · PDF · photo extract Structured record Straight through Human review Straight through low confidence only
Confidence decides the route. You set the review rate you can actually staff.
Mode 03

Enable

Teaching the people who have to live with it. What to use and when not to. How to structure work so it's repeatable. Which decisions are worth automating and which need a person.

For an owner or an executive who wants to work this way personally, this is direct coaching — often the highest-leverage work we do, and occasionally the only thing a client needs from us.

Engagement model

Small team. Forward-deployed.

We work inside your operation, on your systems, alongside the people who do the work. You own the system, the code, the credentials, and the roadmap from day one.

Limited engagements.

Tell us what's breaking. If the answer is a product you should just go buy, we'll tell you that on the first call and point you at it.

Request an engagement
03Field notes

Notes from the field.

Four essays

Short essays on why transformation stalls in operating businesses, and what actually holds up. No vendor talk. No model names.

Recognize your operation?

These come from inside businesses like yours. If one of them described your Tuesday, we should talk.

Request an engagement
← Field notes
Note 01August 20267 min

The high priests of complexity.

There is a great deal of money to be made right now in making this look harder than it is.

Not by lying. Nobody has to lie. It's enough to answer a simple question with a complicated answer, to use six words the client doesn't know, and to let the resulting silence do the work. The client concludes that the problem is beyond them — which is exactly the conclusion that makes them hire someone, keep paying them, and never quite feel able to ask what they're paying for.

This is not new. Every technology wave produces a priesthood: people whose position depends on the thing staying mysterious. What's different now is the speed. The vocabulary changes every few months, which means nobody outside the field can tell the difference between a person who's keeping up and a person who's keeping you confused.

The incentive is structural, not moral.

Most people doing this aren't cynical. They're responding to incentives that all point the same way.

A firm billing for implementation makes more from a custom build than from telling you to buy a product for $40 a month. A firm billing hourly makes more from a system that requires them to maintain it. An internal champion who has staked a promotion on an AI initiative cannot come back and say the answer was a policy change. And a vendor whose product is a hammer will, reliably and sincerely, diagnose a nail.

None of that requires bad faith. It just requires nobody in the room being paid to argue for less.

What "less" often looks like.

In operating businesses — the ones with real revenue, real customers, and processes that grew rather than being designed — the highest-return interventions are frequently unglamorous:

  • A product that already exists. Mature, supported, tested by thousands of companies, with a support line you can call at 2am. Custom software you must maintain forever is a liability you've been sold as an asset.
  • An incentive change. If the weekly report is always late, look at whether anyone is measured on it before you automate its production. Some of the highest-return changes we've recommended cost nothing to implement.
  • Deterministic software. Rules engines, scrapers, scheduled jobs, validation logic. Twenty-year-old technology that runs the same way every time, costs nearly nothing, and never hallucinates. A surprising share of what gets pitched as AI is a job for a scheduled task and a regular expression.
  • Writing the process down. Often the actual deliverable. You cannot automate a process nobody has specified, and the act of specifying it frequently reveals that three of the seven steps exist because of a decision made in 2011 that nobody has revisited.

None of this means the sophisticated version is never right. Sometimes it is, and when it is, it's worth doing properly. The point is that the decision should be made on the merits, by someone who would genuinely have been willing to recommend the cheap answer.

Four questions that expose it.

You don't need technical knowledge to run these. Ask them in the first meeting.

"What would it take to remove this piece?" Point at any component. A person who designed the system can tell you what depends on it and what would break. A person who added it because it's what everyone builds will give you a reason that sounds like a brochure.

"What's the cheapest thing that would get us 70% of this?" Everyone has an answer. The question is whether they'll say it out loud, and whether they treat it as a serious option or as a strawman to dismiss.

"What happens when this is wrong?" Every one of these systems is wrong sometimes. Someone who has actually run one in production will describe error rates, review queues, and what the person on the other end does about it. Someone who hasn't will tell you it's very accurate.

"Explain that again, without the vocabulary." This is the one that matters most. Anything real in this field can be explained plainly to a smart person who works somewhere else. If it can't be, one of two things is true: they don't understand it well enough to simplify it, or they don't want to. Both should end the meeting.

The test that runs in the other direction.

The uncomfortable corollary: restraint is only credible from someone visibly capable of the complicated version. "You don't need this" from someone who couldn't build it anyway is not advice, it's a limitation being described as a philosophy.

So ask for both. Ask what the ambitious version would look like, in detail, and see whether they can describe it precisely — the failure modes, the cost, what would have to be true for it to be worth it. Then ask what they'd actually recommend, and why it's different. The gap between those two answers is where you'll find out whether you're talking to an advisor or a priest.

← Field notes
Note 02August 20266 min

Buy, build, or change the incentive.

A crew of field technicians won't file expense reports. The mobile app has been deployed twice. There have been two memos and one meeting. Receipts still arrive as text-message photos sent to the office manager, who types them into the accounting system on Friday afternoons.

There are three real answers to this, and almost every firm you talk to will only offer you one of them.

Answer one: buy the tool and force the behavior.

Deploy a proper expense platform. It's mature, it's supported, it handles receipt capture, approvals, card reconciliation, and the accounting export. It costs a few dollars per user per month and it will work correctly forever without you maintaining anything.

The catch is in the second half of the sentence: and force the behavior. This answer requires fourteen field technicians to change how they've done something for years. That's not a software project, it's a management project, and it succeeds or fails on whether the owner will actually enforce it — including with the two best technicians, who are the least replaceable and the most resistant.

Answer two: build something that meets them where they are.

Keep the texting. Build a system that receives the messages, reads the receipts, extracts vendor, amount, date, and job code, matches them against the card feed, and posts coded transactions into the accounting system — flagging anything it isn't sure about for the office manager to resolve in a queue.

Nobody has to change anything. Adoption is guaranteed, because adoption already happened years ago. The cost is that you now own a system: it needs maintenance, it will break when a carrier changes a message format, and somebody has to be responsible for it in two years.

Answer three: change the incentive and buy nothing.

Ask why the receipts are late rather than how to process them faster. Sometimes the answer is that reimbursement takes three weeks, so nobody is in a hurry. Sometimes it's that the job code isn't knowable at the moment of purchase, so the technician is being asked to supply information they don't have yet. Sometimes the process is simply nobody's job.

Change the reimbursement cycle to weekly, or move coding to the person who actually knows the job code, and the problem can quietly disappear without any software at all. This answer is the one least likely to be offered by someone selling software, and it is more often correct than that industry's revenue would suggest.

The variable that decides it.

Notice that the three answers are not distinguished by technical merit. They're distinguished by one question, which is a question about your business rather than about technology:

How much behavior change can this team absorb right now?

That number is real, it's finite, and it's usually much smaller than owners think. It's also a budget you're spending elsewhere — if you've just changed the pay structure, or brought in a new ops manager, or moved to a new scheduling system, there may be nothing left this quarter.

A proven tool that your people won't use is worse than a rough system built around what they already do, because the first one costs you money and credibility while the second one at least works. Equally, a custom build you'll maintain forever is a bad trade when a mature product would have been accepted without complaint.

Say it out loud.

The failure mode isn't picking wrong. It's picking without noticing you were choosing.

Most implementations pick answer one or answer two silently, based on what the vendor happens to sell, and the behavior-change cost never appears in the business case. Then adoption fails, and everyone concludes the technology didn't work — when what actually happened is that a management decision got made by default, by someone who wasn't in a position to make it.

So we write it down. For every candidate on the roadmap: what we'd buy, what we'd build, what we'd change instead, what each costs, and how much behavior change each one asks of your people. Then the owner decides, on the record, with the tradeoff visible. That's a decision. Everything else is a drift.

← Field notes
Note 03August 20266 min

Why your pilot produced a slide deck.

Somewhere in your industry, a company ran a pilot last year. It went well. The demo impressed everyone. The consultants presented a readout with a projected ROI. And today, nobody at that company uses it.

This is the normal outcome. Not because the technology failed — it usually worked fine in the demo — but because of four structural decisions made before the pilot started, each of which guaranteed the result.

Pilots are scoped to demonstrate, not to carry load.

A pilot's success condition is a good demo. So it gets built on the twenty cleanest examples: the well-scanned invoices, the standard work orders, the accounts with complete records. The messy 30% — the faxed statement, the handwritten field note, the account that exists twice under two spellings — is excluded as "out of scope for the pilot."

But the messy 30% is the job. That's where the handling time is, where the errors are, and where the person you're trying to help spends their afternoon. A system that handles the clean 70% doesn't reduce anyone's workload; it adds a triage step. The person still has to look at everything to decide what the system can handle. You've automated the easy part of their day and left them the worst of it.

Nobody instruments a pilot.

Production systems get dashboards. Pilots get anecdotes. When nobody logs every run, captures every correction, and publishes an error rate, the decision at the end has no evidence to rest on. The champion says it went great; the skeptic says it made a mistake once; both are working from memory. The safest executive decision in that meeting is to not decide — which is what "let's extend the pilot" means.

Nobody was brought along.

The people who would have to live with the system saw it in a demo and got a training deck by email. Nobody sat next to them during a real week of work, watched where it got in the way, and adjusted it. So at the first inconvenient moment they went back to the spreadsheet, which never fails and never surprises them.

Adoption is not a rollout phase that happens after the build. It's most of the build, and it's the part that gets cut when the timeline slips.

Nobody owns it.

Ask who owns the pilot and you'll get a committee: the consultant built it, IT hosts it, an ops person sponsors it. Ask who fixes it when the invoice format changes in January, and the room goes quiet. A system without a named owner decays at the speed of its most fragile integration — usually about six months.

The alternative is not a bigger pilot.

It's a smaller production system. One workflow, chosen because it's boring, high-volume, and reversible. Deployed into real work on day one, with a human reviewing every output. Instrumented, so the error rate is a number on a dashboard rather than an argument in a meeting. Shaped around how the team already works. And owned — by a named person inside your business, trained on a written runbook, before the builders leave.

The uncomfortable summary: a pilot is a way of not deciding. If you're not willing to put something into production with a human safety net, you haven't found the right workflow yet — and that's a selection problem, not a technology problem.

← Field notes
Note 04August 20267 min

What accuracy actually means for document extraction.

A vendor tells you their extraction is "98% accurate." That sounds like an A+. Here is the arithmetic they're hoping you won't do.

Field-level vs. document-level.

A carrier statement, an invoice, a work order — each is not one answer but many. A typical commercial invoice might carry 30 fields you care about: vendor, date, PO number, line items, quantities, rates, tax, total. "98% accurate" almost always means per field. The number you actually live with is per document — the odds that a document comes through with every field right:

field accuracy 0.98 fields per document 30 document accuracy 0.98^30 ≈ 0.545 → nearly half of documents contain at least one wrong field

At 98% field accuracy, roughly 45% of your documents have at least one error in them somewhere. If a document with one wrong field is a document a human must fix — and in billing, it is — your "98% accurate" system requires review of every document, and you've saved the typing but kept all of the checking. Sometimes that's still a good trade. But it's a very different promise than "98%."

Confidence is the tool that changes the economics.

Good extraction doesn't just output a value; it outputs a value and a confidence. That turns one queue into two:

High-confidence output flows straight through, spot-checked by sampling. Low-confidence output goes to a human review queue, presented next to the source document so the check takes seconds rather than minutes.

Now the operating question stops being "how accurate is it?" and becomes "what review rate can I afford?" — a question about your business, not about the technology. If a person can clear 60 flagged documents an hour and you process 4,000 documents a month, a 20% review rate costs about 13 hours of skilled review per month against, typically, several hundred hours of manual entry displaced. You set the confidence threshold to hit the review rate you can staff — and you tighten it for fields where an error is expensive (rates, totals) and relax it where an error is cheap (a memo line).

Measure on your documents. Only yours.

Any accuracy number measured on someone else's documents is marketing. Scan quality, form variety, handwriting, seasonal volume spikes, the carrier that redesigns its statement every January — these are properties of your paper, and they dominate the result. The honest procedure: take a few hundred of your real documents, including the ugly ones. Have the system extract them. Have a person mark every field right or wrong. That number — measured on your documents, at the field level and the document level, before anything is trusted — is the only accuracy figure that should appear in a business case.

If a vendor resists measuring on your documents before quoting a number, that tells you the number.

04Inquire

Limited engagements.

We take a small number of engagements at a time, because the work is hands-on and done inside your operation. We're direct about fit — if what you describe is better solved by a product you should just go buy, a process change, or a different firm, we'll tell you that and point you at it.

Write to us
engage@directresult.co

The more specific you are, the faster we can tell you whether there's something here. Worth including:

  • What the business does, and roughly what size.
  • The work you want to stop doing by hand — the documents, the systems, and who does it today.
  • What you've already tried, if anything, and where it stalled.

You'll hear back within two business days. If it's a fit, the next step is a short call to scope the survey.