RAG Development Services

    Retrieval-augmented generation over your real document set, with citations, permission filtering and an evaluation harness that proves the answers hold up.

    RAG, or retrieval-augmented generation, is the technique of retrieving relevant passages from your own documents and giving them to a language model so its answers are grounded in your content rather than its training data. AMT builds enterprise RAG systems with hybrid retrieval, reranking, source citation, permission-aware access and a scored evaluation harness. We are ISO 27001:2013 certified.

    // the_problem

    When RAG gives a wrong answer, the model is rarely why

    The retrieval is. Chunks split mid-clause. A policy document and its superseded version both sitting in the index. Semantic search that cannot match an exact product code. No way to tell which source an answer came from, so nobody can verify it and nobody trusts it.

    And almost nobody builds the evaluation harness. Without one there is no way to know whether a change improved the system or quietly broke it, which means the safest thing to do with a working RAG system becomes nothing at all.

    // the_readiness_test

    Is your document set ready for RAG?

    The single biggest predictor of what a RAG project costs, and whether it works, is the state of the documents. Check yours against this before anyone quotes you.

    QuestionWhat a bad answer means
    Is there one current version of each document?If superseded versions are still in the folder, retrieval will surface them alongside the correct ones and you will not know which answer came from where.
    Are the documents in consistent formats?A mix of PDFs, scanned images, spreadsheets and email threads roughly doubles ingestion effort.
    Does someone own keeping them current?A corpus with no owner degrades, and so does the system built on it. This is an ownership question, not a technical one.
    Do exact identifiers matter, such as product codes, clause numbers or part numbers?If yes, semantic search alone will miss them and you need hybrid retrieval. Many implementations skip this and nobody notices until a customer does.
    Do different people have different access rights to these documents?If yes, retrieval must be permission-filtered per request. Retrofitting that later is painful and it is a security incident waiting to be discovered.

    Score your own corpus first

    The five questions with the scoring guide, on one page. Run it before anyone quotes you, because document condition drives the price far more than document count does.

    No newsletter, no sequence. One email with the file attached.

    // what_we_build

    What we build

    Document ingestion
    Parsing and chunking tuned per document type, not one global chunk size.
    Hybrid retrieval
    Semantic plus keyword, so exact identifiers still match.
    Reranking
    A second pass that fixes what first-stage retrieval gets wrong.
    Citation and traceability
    Every answer traceable to a source passage.
    Evaluation harness
    A scored test set run on every release.
    Freshness pipelines
    Re-indexing as documents change, with versioning.
    Access control
    Retrieval that respects who is allowed to see what.

    // proof

    Proof

    Enterprise knowledge base RAG pipeline

    Retrieval-augmented generation across an internal knowledge corpus using LLMs and vector search, with citation back to source passages.

    Telecom conversational AI agent

    RAG over policy documents, FAQs and product catalogue combined with live API calls. 70% of queries resolved without a human agent, under 3 seconds average response.

    // cost

    What RAG costs

    Document condition drives this more than document count. A clean, consistent, well-owned corpus of fifty thousand documents is cheaper to build on than five thousand documents spread across fifteen years of formats, scanned images and inconsistent naming.

    RAG implementation effort depends on the quality and volume of source data, ingestion requirements, access controls, integrations and the level of accuracy and governance required in production.

    ScopeTypical complexityWhat drives the effort
    RAG on a defined, clean corpusLowData volume, document formats, ingestion and indexing complexity, retrieval quality and evaluation requirements.
    RAG with permission filteringMediumAdds user- or role-based access controls, permission-aware retrieval and testing against the organisation's actual security model.
    RAG within a wider AI agent solutionSee AI AgentsRetrieval becomes one component of a broader solution involving agents, enterprise integrations, workflows, guardrails and production monitoring.

    // the_honest_section

    Who this is not for

    If your documents do not exist in any system we can reach, the first project is a document project. We will say so, and it is cheaper to hear that now.

    If the answers you need are calculations rather than passages, RAG is the wrong tool. A query against a database will be faster, cheaper and correct every time.

    If nobody owns keeping the corpus current, the system will degrade and we will both be disappointed in a year. Name the owner before the project starts.

    // faqs

    Frequently asked questions

    What is RAG?

    Retrieval-augmented generation. The system retrieves relevant passages from your own documents and gives them to a language model, so answers are grounded in your content and can be traced back to a source.

    Why does our RAG system give wrong answers?

    Usually retrieval rather than the model. Common causes: chunks split mid-clause, superseded document versions still indexed, semantic search missing exact identifiers, and no reranking. The corpus readiness questions above will point at which one applies.

    RAG or fine-tuning?

    RAG when the knowledge changes and has to be current and citable. Fine-tuning when you need a consistent behaviour, format or tone. They solve different problems and are frequently used together.

    How do you evaluate a RAG system?

    A scored test set of real questions with known correct answers, run on every change, measuring retrieval quality separately from answer quality. Separating the two is what makes a failure diagnosable.

    How do you handle documents different people are allowed to see?

    Retrieval is permission-filtered per request against your existing access model, so the system cannot surface a passage the person asking is not entitled to read. This is designed in from the start.

    How many documents do we need?

    Fewer than people expect. Condition matters far more than count. A few thousand clean, current, well-structured documents outperform hundreds of thousands of mixed ones.

    What does RAG development cost?

    See the table above. Document condition is the main driver.

    Can you use our existing vector database?

    Yes. We work with what you have where it fits, and tell you plainly when it does not.

    Tell us about your documents

    Thirty minutes with an engineer who has built retrieval over real enterprise corpora. Describe what you have and we will tell you what it takes to make it answerable.

    We reply within one business day. Your first conversation will be with a senior technical partner, not a salesperson.

    We do not share your details, we do not sell lists, and we will not add you to a newsletter you did not ask for. Read our privacy policy.

    Not ready to talk? Read how we deliver, or see what software development costs. No form.