Building an AI assistant that respects document permissions

Connect a model to SharePoint or Confluence and it can answer from files the person asking was never allowed to open. Here is how to carry permissions through to every answer.

Datagist3 min read

A search box only shows people the files they are allowed to open. An AI assistant built on the same files does not always behave that way.

The usual design, often called retrieval-augmented generation, copies documents into a search index, finds the passages most relevant to a question, and asks a model to answer from them. If the index does not know who may read each passage, the assistant can answer from the salary file or the board pack for anyone who asks. The person never opens the file. The assistant simply tells them what is in it.

This is one of the most common reasons an internal assistant is held back from launch. It is also a solvable problem, if permissions are designed in from the start.

Carry permissions into the index

When a document is loaded into the index, record who can read it: the groups and people allowed, and any who are explicitly denied. Read these from the source system, such as SharePoint, OneDrive, or Confluence, so the assistant follows the same rules people already manage there.

Three rules keep this safe. The first is that the default is nobody: a file with no matching permission rule is not indexed at all, because it is better to leave a document out than to guess who may see it. The second is that passages inherit their document’s permissions. Documents are split into passages for search, and splitting must never widen who can see something. The third is that a denial wins. If a person belongs to one group that is allowed and another that is denied, the denial applies.

Filter before ranking

When someone asks a question, filter the index down to what that person can read before the search ranks any results. Passages they cannot read are then never scored, never returned, and never reach the model.

Filtering after the search is weaker. By the time results are filtered, the model or the logs may already have seen restricted content, and some designs leak it through summaries or suggested follow-up questions.

Then check again. After retrieval, re-check every result against the person’s permissions and raise an alert if any fails. In a correct system the alert never fires. If it does, you want to know immediately.

Treat document content as untrusted

Documents can contain text written to steer an AI system, such as a hidden line in a supplier FAQ telling the assistant to ignore its rules. This is known as prompt injection, and it needs defending against in layers.

At ingestion, screen documents and leave suspicious passages out, while keeping the rest of the document searchable. When passages are handed to the model, mark them as material to answer from, never as instructions to follow. And before every release, test the assistant against injection attempts. No single layer catches everything, which is why all three are needed.

Keep a record and set limits

Log every request: who asked, what was retrieved, and what was answered. Cite the source documents in every answer, so that people can check them and reviewers can trace any answer back to where it came from. Set daily limits on requests and cost too, so a runaway script or a misuse cannot run up a large bill.

Plan for stale permissions

Permissions in the index are only as current as the last sync. If someone loses access to a folder at 10am and the next sync runs overnight, the assistant may still answer from that folder until then. Agree a sync schedule that matches how sensitive the content is, and sync the most sensitive sources most often.

Test it the way an attacker would

Before launch, build a test set from your own content that includes at least one permission test for each restricted area: a question that only authorized people should get an answer to. Run each one as an authorized user and as an unauthorized one. The first should get the answer with a citation. The second should be told the assistant could not find it in the documents they have access to.

Run these tests on every change to the assistant, alongside the tests for accuracy and injection, so that a change made for one reason never weakens the permissions by accident.


Our Secure RAG Starter implements this design on Azure AI Search and Azure OpenAI, with connectors for SharePoint, OneDrive, and Confluence. If an internal assistant is waiting on a security review, tell us about it and we will tell you what the first step would be.

Working on something like this?

Tell us the business decision you want to improve. We will tell you what the first step would be and whether we are the right fit.

Start a conversation