Kuroko SystemsConsult

Insights

Past the chatbot: where language models actually save work

5 min read

A chat bubble in the corner is the most visible thing you can do with a language model, and one of the least useful. It faces customers who mostly want a fast answer from a person, and it is measured in satisfaction that is hard to measure.

The work actually being saved is backstage: in a search that takes an employee twenty minutes, in retyping data from a document into a system, and in the preparation that comes before every decision. There the measurement is simple — how long did this take before, and how long does it take now.

The problem is not missing information, it is unfindable information

In an average organisation the information is spread across folders in a drive, contracts in PDFs, customer records in a CRM, and long email threads. None of those sources is missing — you just cannot ask them a question.

That is what grounded retrieval solves: rather than retraining a model on your material, you connect it, under control, to the stores that already exist. The model searches, reads, and answers out of what it found.

The difference in practice: someone who gets a technical question mid-call does not stop, search a forty-page specification, and come back an hour later. They type the question, get an answer, and see a link beside it to the source it came from.

That link is not decoration. An answer without a source is an answer nobody can check, and a system nobody can check is one nobody should trust with a decision that costs money.

From documents into systems, without typing

A large share of manual work in any organisation is moving data from one format into another: a purchase order that arrived as a file, a lead written into the body of a long email, a service request that came in as a message. Until now somebody typed those into a system.

Language models are very good at exactly this — pulling structured fields out of free text. An automation catches the message, the model extracts the name, the product, the quantity and the urgency, and the data lands in the system as a clean row.

What needs saying and is not said often enough: this does not reach zero errors. What changes is the kind of error — typing mistakes are traded for extraction mistakes, which are rarer but quieter, because nobody watches them happen.

So extraction only works with a safety net: the model returns a confidence along with the answer, anything below the line goes to a person, and every field is stored next to the source it came from so it can be checked. With that, the time saved is real. Without it, you have built a machine that produces errors faster.

An agent that prepares, a person who approves

The next step is moving from a model that reads and summarises to one that carries out a sequence of actions. That sounds risky, and it is — unless the decision itself stays with a person.

Here is what a working process looks like, on a refund request for a faulty product:

  1. Identify — the system recognises the message as a refund request rather than a question or a general complaint.
  2. Investigate — the agent checks against the store when the item was bought, whether it is under warranty, and what the customer's history is.
  3. Prepare — it drafts a reply and creates a draft credit note in the accounting system. Both are drafts, and neither goes out.
  4. Approve — the shift manager sees all of it on one screen, including what the agent checked and on what basis, and approves or rejects.

Where the data actually sits

Once a model is running over contracts, leads, and financial figures, the question of where the data travels stops being technical and becomes legal. And there is a distinction here worth not blurring.

On the major providers' business tiers your data is not used to train models — that is contractual rather than a promise. But it does pass through their infrastructure. That is a level of risk most organisations accept quietly, the same way they accept it for their email and their cloud storage.

If the material is sensitive enough that even that transit is a problem — regulation, a client who forbids it, or data that must not leave the network — then the answer is an open model running on infrastructure you control. It costs more to stand up and more to maintain, and it is the only arrangement where “the data never leaves us” is precisely true.

Choosing between the two is a business decision rather than a technical one, and it is worth making before the system is built rather than after.

Knowing that it works, and what it costs

A model-based system does not fail loudly. It starts answering slightly less well, and if nobody is measuring, you find out from a customer.

So an internal tool needs a set of test cases with known good answers, run on every change. That is what lets you say a model swap or a reworded prompt is an improvement rather than a regression, instead of forming an impression.

And cost has a way of climbing quietly. A larger model, a longer prompt, or a call running in a loop all look identical from outside and cost several times as much. We pick the cheapest model that passes the tests rather than the strongest one available, and put a ceiling and monitoring on it — so a surprise arrives as an alert rather than as an invoice.

An internal AI tool is not measured by impression or by a vague notion of customer experience. It is measured in hours handed back to the people who were spending them, in a response time that got shorter, and in how much less often somebody typed a figure that was already written down somewhere else.

Which is also why the work suits us: these tools are at their best when nobody notices them.

Common questions

Where do we start if we have nothing?

With one repetitive process that irritates somebody. Not with the technology. Look at what an employee does several times a day in exactly the same way — triaging enquiries, pulling details out of a document, searching the same files. The first tool should be small enough that you can abandon it if it does not work.

Does this replace staff?

With internal tools, usually not — they remove the part of a job nobody wants to do rather than the job. Anyone looking for headcount cuts tends to find they have built a system that needs supervising rather than one that saves it. Anyone looking for their team to do less clerical work and more real work gets that.

Do we need to train our own model?

Almost never. Training is expensive, goes stale, and needs a volume of data most organisations do not have. Retrieval over your own documents gets the same result for a fraction of the cost, and updates the moment a document changes rather than requiring another training run.

How long does a first internal tool take?

By scope, and less than most people assume — because a first tool should be small. Most of the time does not go on the model but on access to the data: where it is, who is allowed to see what, and what state it is in. You get the estimate together with the scope.

What we do about this

  • AI Systems & AutomationAssistants, document handling, and search across your own material — built on models you already pay for, with a person wherever the decision costs money.
  • Business DashboardsSeparate data sources brought into one view, with live figures where they matter, alerts when something moves out of range, and access by role.

Working on something like this?

Tell us the shape of the problem and you get back scope, risks, and a recommended way of working — usually within one business day.

Consult

All insights