AI-powered education platform
Personalised learning paths and skill assessments generated per learner, with output constrained to the curriculum rather than left open-ended.
LLM applications that stay reliable at production volume, with the run rate modelled before the build starts.
Generative AI development is the building of applications on large language models: knowledge assistants grounded in your content, structured extraction from documents, constrained content and code generation, and fine-tuned models where prompting alone will not hold. AMT builds these on OpenAI, Anthropic and open models, provider-agnostic where the architecture allows it, under ISO 27001:2013 information security practices.
// the_problem
Getting a language model to produce something impressive takes an afternoon. Getting it to produce the right thing, in the right format, every time, across inputs nobody anticipated, at a cost that still works when ten thousand people use it, is the actual project.
Three things decide whether a GenAI application survives contact with users: whether it is grounded in something real, whether its output is constrained to a shape your systems can consume, and whether anybody is scoring it.
// what_we_build
// method
// cost
Build cost is only half the number. The running cost is the half that surprises people: token spend at production volume, which is routinely many times what the pilot suggested, because the pilot was ten users and production is ten thousand.
Four things drive the run rate, and only one of them is the model price.
| Driver | Why it moves the number |
|---|---|
| Context size per call | The largest lever by far. A system that stuffs twenty documents into every prompt costs many times one that retrieves three good ones. |
| Calls per user per session | Chained calls and retries multiply quietly. A three-step chain is three times the spend of the one-step version nobody remembers approving. |
| Output length | Cheaper to generate a structured answer than an essay, and usually more useful too. |
| Caching and routing | Sending easy queries to a small model and caching repeated ones is unglamorous and routinely the difference between viable and not. |
We model the run rate before the build starts, not after the first invoice. If the economics do not work at your expected volume, that is worth knowing in week one.
Work out your own run rate
Put in your expected users, queries per user and context size, and this returns a monthly token cost. It is the number that surprises people in month two, and it is better to see it in week one.
// proof
Personalised learning paths and skill assessments generated per learner, with output constrained to the curriculum rather than left open-ended.
Retrieval-augmented generation across an internal corpus, every answer traceable to a source passage.
// the_honest_section
If the output has to be correct every single time with no human in the loop, such as a medical dosage, a legal filing or a financial calculation, generative AI on its own is not the right tool. Any firm telling you otherwise is selling.
If your use case is well served by a template or a rules engine, that will be cheaper, faster and more predictable. We will tell you when that is the case.
If the run rate does not work at your expected volume, we would rather establish that in week one than build something you have to switch off in month six.
// faqs
Building applications on large language models: assistants, extraction, constrained generation and fine-tuned models, engineered so the output is reliable enough for real use rather than impressive in a demo.
You ground it in real sources, constrain the output to a defined schema, define refusal behaviour for questions it should not answer, and score it against a test set on every release. Hallucination is not eliminated; it is bounded, detected and made visible with citations.
It depends on your accuracy needs, latency budget, data residency requirements and volume. We build provider-agnostic where the architecture allows, so the choice stays reversible when pricing or availability changes.
Often yes, particularly where data residency or run rate matters. The trade is usually more engineering effort for lower marginal cost, which pays back above a certain volume.
See the drivers above. Context size per call is the largest lever, and it is the one most implementations never tune.
A scored test set of real inputs with known good outputs, run automatically on every change, with accuracy, cost and latency tracked per release.
No, unless you ask us to and the contract defines the boundary. Your data is never used to train anything outside your own engagement.
Thirty minutes with an engineer. Bring your expected volume and we will tell you whether the unit economics work, and what to change if they do not.
Not ready to talk? Read how we deliver, or see what software development costs. No form.