Key Takeaways
- Run the ai vendor security questions at your vendor before the security review runs them at you. The weeks you lose are almost never spent on the answer. They are spent finding out you didn't have one.
- Ask when the vendor decides what a person is allowed to see: before it goes looking, or after it has found things. Almost everything else on the list is downstream of that answer.
- Don't accept a certificate as the answer to a question it was never asked. A clean audit report describes the vendor's own perimeter, not what your contractor can get the assistant to say.
- Ask about all five places content gets out: what the assistant finds, what it counts, what it writes, what it remembers and what it logs. A vendor who answers the first and waves at the rest has a demo, not an answer.
- Watch the failure tests instead of hearing about them. Someone with no access at all, someone just switched off, a member of the public: three empty results is a pass.
- Avoid any tool that keeps its own separate list of your people. If removing a leaver takes two actions in two places, one day it will take one, and you will find out after they have gone.
- Expect to be judged on what happens when it goes wrong. One design fails by showing nothing, which you hear about within the hour. The other fails by showing everything, which you hear about months later, from someone forwarding an email.
- Own the part that is yours. What the assistant returns is only as correct as who-can-see-what in your own systems, so put group membership on a review cycle before security asks whether you have one.
The security review is scheduled for thirty minutes. It takes eleven.
Carmen has done her homework. She has brought the vendor's SOC 2 Type II report, their data processing agreement, their penetration test summary and a one-page diagram. Her security architect flips past all of it. Then he asks the question that ends the meeting: when someone asks the assistant something, how do you know it can't tell them something they aren't allowed to see?
Carmen doesn't know. The vendor's sales engineer said "we respect your existing permissions," and she wrote that down, and at the time it sounded like an answer. It isn't one. It is a description of an intent. The review is rescheduled three weeks out, the pilot slips a quarter, and she goes back to the vendor with a question she can't evaluate the answer to.
The other version is worse.
It happens after you have already won. The review passes, you launch, and four months later a support manager forwards an answer to a colleague who replies "wait, how do you have that?" The assistant had quoted a churn-risk memo written for the CRO's staff meeting, including the line about the executive sponsor who stopped replying to emails. Nobody attacked anything. Nobody guessed a URL or bypassed a login. It found the most relevant material in the company and explained it clearly to the person who asked, which is exactly what you bought it to do.
Both versions come out of the same gap, and it is a gap you can close on your own schedule instead of the reviewer's. You don't need to become technical. You need to know which questions separate a vendor who has solved this from one who has described it, and to ask them while you still have the leverage of not having signed.
A clean SOC 2 report doesn't answer the question you're actually asking
Here is the claim, and it is meant to be arguable: a SOC 2 Type II report tells you almost nothing about whether an AI vendor will show one of your users content they aren't allowed to see. Not "less than you'd hope." Almost nothing.
The reason is structural. SOC 2 is an attestation against the AICPA's Trust Services Criteria, and the auditor tests the controls that management describes. If a vendor says "access to customer data is restricted to authorised personnel," the auditor samples that control and tests whether it operated. That control is about the vendor's staff reaching your material. It is not about whether your marketing contractor's question to the assistant comes back with a paragraph from an unreleased pricing document, because nobody wrote that control, because the criteria don't ask for it.
You can push back on this, and a good security person will. The logical access criteria are broad enough to cover what one of your users can get out of the product, and a rigorous auditor with a well-scoped description will test it. That's true. It is also true that scope is chosen by the vendor, that assistants of this kind are new enough that most descriptions predate them, and that you cannot tell from the report cover which situation you are in. A certification tells you a company has a functioning control process. It doesn't tell you which controls they pointed it at.
The newer frameworks are closer to the problem and still not on it. ISO/IEC 42001 governs how an organisation manages AI systems, which is genuinely useful and still a management-system standard. NIST's AI Risk Management Framework is voluntary guidance for identifying risk, not a test you pass. OWASP's Top 10 for LLM Applications comes closest, naming sensitive information disclosure and excessive agency as real categories, and it is a threat list rather than something a vendor gets audited against.
So the compliance layer is real and it covers the vendor's perimeter. The questions below cover yours.
Two products that look identical in a demo
There are two ways to build an assistant that respects who is allowed to see what, and in a forty-minute demo they are indistinguishable. Both give the right person the right answer. The difference only shows up on the day something breaks, which is also the day it becomes your problem.
The first kind looks everywhere, then tidies up. It finds the best material in the company, works out which of it this person isn't allowed to see, and takes that away before anything reaches the screen. Whatever survives gets shown.
The second kind never goes looking outside what the person is allowed to see. Material beyond that boundary is never a candidate. It is never found, never summarised and never quoted, because it never enters the process at all.
The claim worth carrying into your review is this. Tidying up afterwards is not access control. It is redaction, and redaction is a decision about what to display. A product that finds everything and then hides most of it is one mistake away from showing it, and these products grow new places to display things faster than anyone can check them.
There is a third answer you will hear, and it is the most common of all: permissions aren't part of the picture at all, and instead you are asked to designate a "safe" set of content for the assistant to use. That is a workaround rather than a model, and it collapses the moment your real audiences have more than a couple of tiers.
Your security team will push back, and they have a point
Expect this, because it is the strongest objection and a fair one. Removing results before the user sees them is, formally, access control: the check happens before anything reaches the person. Search products shipped exactly that design for fifteen years without notable incident. Three things changed, and they compound. It is worth being able to say them out loud, because the reviewer who raises the objection is usually the person you most need on your side.
The first is how many places an answer now appears. A search box used to have one output: a list of links. Your assistant has a dozen. The results list, the answer panel, the chat thread, the type-ahead, the counts in the sidebar, the related-content rail, the export, the Slack preview, the weekly digest, the dashboard your ops lead built. Tidying up has to work correctly on every one of them, forever, including the two your vendor ships next quarter.
The second is that the thing being removed changed. You can withhold a document. You cannot withhold a sentence the assistant wrote after reading that document. Once material has been read, it has been turned into prose, and prose carries no record of who was allowed to see it.
The third should settle it, and it is the asymmetry to put in front of your reviewer in exactly these terms. The two designs fail in opposite directions. The one that never looks outside your boundary fails by showing nothing. That is an outage: loud, immediate, and on your desk within the hour, because the people it happens to are your own staff and they will tell you at once. The one that tidies up afterwards fails by showing everything, because the step that was supposed to remove things is the step that broke. That is a breach, and it is silent. Nobody is inconvenienced by it, so nobody reports it. You hear about it months later, from someone forwarding an email, and the question in the room is no longer whether it happened but how long it had been happening.
One design degrades into a bad afternoon. The other degrades into a disclosure conversation you have with a lawyer present. So the question is not whether tidying up can be correct. It can. The question is what happens on the day it isn't, who finds out first, and whether that person works for you.
Why this got harder the moment you added AI
The old design survived because a person was the last step. Ten blue links went to a human who read the titles, clicked one or two, and moved on. If the filter slipped, one person saw one file name.
An assistant removes that buffer. It reads everything that comes back, in full, before anyone decides what to put on screen. If something is pulled back and then dropped before display, it still shaped the answer: the influence survives the deletion. And it is a far better reader than your support manager was. It doesn't skim, it doesn't get bored, and it will happily join up forty documents no employee would ever have opened in a row.
Your security team is not inventing this concern. OWASP moved Sensitive Information Disclosure from sixth place to second in its 2025 Top 10 for LLM Applications, and its guidance is blunt about the most popular reassurance vendors offer: instructions telling the model what not to reveal "may not always be honored and could be bypassed via prompt injection or other methods." Telling the assistant to behave is not a boundary. If that is a vendor's answer to your permissions question, you have your answer.
The exposure underneath it is not new, which is the part worth telling your CFO. Varonis analysed roughly 10 billion files across 1,000 real environments and found that 99% of organisations have sensitive data that AI can surface, 66% have cloud data exposed to anonymous users, and 88% carry stale but still-enabled ghost accounts. None of that was created by AI. It was created by a decade of over-sharing, and it sat quietly because finding it required knowing what to look for. Asking a question in plain English removed that requirement. "What are we worried about with Northwind" does not need a file name.
Which means your project is now carrying a risk your company created years before you arrived. That is unfair, and it is still yours to answer for. The way out is to be the person who arrives at the review with the questions already asked.
The question set, in the order that decides things
Thirteen questions, grouped by what they are for. The first five are about where content physically gets out, and if those go badly you can stop, because nothing later rescues them. The next three are about the edges, where systems of this kind actually fail. The last five are the contractual and operational half, where you will live once the thing is running. For each one: what you are really asking, the answer that should worry you, and the answer that should let you move on.
Ask them in an evaluation call rather than sending them as a document, because you are listening for hesitation as much as content. If your own team is building any of this in-house, the same questions work pointed inward, and it is better that you ask them than that your reviewer does.
Part one: the five places content gets out
Nearly every real incident of this kind escapes through one of five routes. Each has a question a vendor either answers precisely or doesn't, and the difference is audible without any technical background at all.
1. What it finds
The assistant returns something the person shouldn't have. Ask: at what point do you decide what this person is allowed to see, before you go looking or after you have found things? You are listening for whether material outside someone's access is ever handled at all. If the answer describes results being removed on the way back, follow it with the question that actually matters to you: everywhere it is shown, or only here? A vendor who has to pause on that has told you something.
The answer that should worry you: "We filter the results before they're shown to the user." "Our assistant only uses the content you sync to us." "We respect your permissions." That last one is the worst, because it sounds like a yes. Follow it with: respect them when, before you look or after?
The answer that should reassure you: what a person is allowed to see is settled before anything is found, so material beyond that is never retrieved, never summarised and never quoted. There is no second pass and no list to subtract from, because there is nothing to subtract.
There is a second reason to care, and it wins the budget conversation rather than the review. Tidying up afterwards wrecks the quality of what is left. If the assistant finds the fifty best matches and then removes forty-two, your agent is handed eight results that were never the best eight for them: the leftovers of a ranking worked out for somebody else. When the boundary is set first, the best answers are the best answers within what that person can actually see, which is the only ranking that ever meant anything to the person asking.
2. What it counts
The document stays hidden but the fact of it doesn't. Counts, totals, "did you mean" suggestions, type-ahead and trending-topic panels are often worked out across everything. A sidebar reading "Legal (14)" tells a contractor there are fourteen legal documents about the Northwind acquisition, which is most of what they wanted to know.
Ask: are counts and suggestions worked out only from what this person can see?
The answer that should reassure you is that the same boundary applies to a number in a sidebar as to a document on screen, and that the vendor says so without being prompted. A pause here, while they work out whether a count is a result, is the finding.
3. What it writes
Something is read, dropped from the screen, and still ends up in the answer. Ask: if the assistant reads something this person shouldn't have, what happens to the answer it has already written using it?
The answer that should worry you: any description of cleaning up the answer afterwards, or of instructing the assistant not to mention certain things. Both put the check after the reading.
The answer that should reassure you: the situation cannot arise, because the material was never read in the first place.
4. What it remembers
Someone else's legitimate answer gets served to you. Saved answers, session summaries, precomputed "popular questions" and shared conversation history all do this, and this route leaks quietly and at scale, because reused answers look like the product being fast.
Ask: can an answer prepared for one person ever be shown to another? Then ask the same question one level up, because these products are shared by many companies at once: can anything prepared for one customer ever reach another?
The answer that should worry you: "every record is tagged with the customer it belongs to." True and insufficient. "We're certified, so that's covered." See the first half of this post.
The answer that should reassure you: the vendor raises reuse before you do, states plainly that a stored answer carries the same boundary as the person it was made for, and can point to the testing that proves it.
5. What it logs
The content escapes through the reporting stack. Search logs, error reports, session recordings, analytics dashboards and the sample material used to tune quality all tend to contain what was retrieved, and they are usually readable by a much wider group than the material itself.
Ask: what do your logs keep, who can read them, and for how long?
What should reassure you is a specific account of what is kept, a named group who can read it, a retention period stated as a number, and an admission of where people at the vendor look at your material and why. What should worry you is a vendor who has never considered that their own quality review is a place your content is stored.
Score all five. A vendor who answers the first well and waves at the other four has a demo, not an access model.
Part two: the three edges, where these systems actually fail
6. What someone with no sign-in gets
Every assistant eventually gets pointed at a place where the person asking isn't signed in: a help centre widget, a pre-login support page, a public documentation site. Every access system also has an edge where it cannot tell who is asking or what they are entitled to, and what happens at that edge is the whole character of the product. There are two possible defaults, and only one was ever a decision anybody made. Either not knowing means no boundary, so the assistant answers from everything, or not knowing means the narrowest possible boundary, so a visitor sees only what was deliberately published and nothing else.
Ask the same question about a partial failure: what happens when the sign-in system is unreachable, or someone's session has expired, or the check on which groups they belong to doesn't come back in time?
The answer that should worry you: "anonymous users get the public knowledge base," with no account of how public gets decided. "We'd never put that on a page outside the login." "That shouldn't happen." Any answer that describes how they intend to deploy it rather than how it behaves.
The answer that should reassure you: anything it cannot positively confirm someone is entitled to returns nothing. A visitor who isn't signed in sees only what a person deliberately published to the public, and the same path handles them as handles your own staff. Ask them to show it live with a restricted document, not on a slide.
7. How fast a change takes effect
Your access rules aren't static. Someone changes teams, a contractor is switched off, a deal-desk document moves out of a restricted space, an employee leaves on a Friday. Ask how long it takes for a change made in your own systems to be honoured by the assistant, and ask specifically about taking access away, which is the direction that matters. Access granted late is an annoyance. Access removed late is the incident. You want the answer as a number, not as "near real time," and you want to know what a person can still get in the meantime.
The answer that should worry you: "permissions come along with the content." "It refreshes overnight." Silence followed by "let me check with engineering," which is fine as long as the answer comes back with a number in it.
The answer that should reassure you: who someone is and what they belong to is established fresh each time they ask, so switching them off takes effect at once, and any part of the picture that is slower to update is named, with its lag stated as a number.
8. What happens when someone leaves
This is the question your reviewer will ask twice, and the fastest way to fail a review is to admit that removing someone from the company does not remove them from the assistant.
Any tool that keeps its own separate list of your people creates that gap. Two lists of who works here means two lists to keep current, and only one of them is anybody's actual job. Removing a leaver then takes two actions in two places, and one day it will take one, and you will find out after they have gone badly. Ask whether access follows the sign-in system your company already uses, or a list kept inside the vendor's product.
Then raise the follow-up yourself, before your security team does, because raising it first is what makes you credible in the room: there is always a short gap between switching someone off and their current session ending. That is true of every signed-in system you already run, so it is not a scandal, but it is a number, and it belongs in the review.
The answer that should worry you: a separate list of users maintained inside the product, removal by hand, or a vendor who says there is no gap at all. They either haven't looked or aren't telling you.
The answer that should reassure you: removing someone once, in the system you already run, removes them here too, with no second action anywhere, and the vendor volunteers the size of the session gap before you ask.
Part three: the paperwork half
9. Whether your content trains anyone's model
Most vendors have a clean answer here, which is exactly why you should read it carefully rather than accept the headline. Three things need separating: whether the vendor learns from your content, whether the AI providers they rely on keep or learn from what passes through them, and whether anyone uses your material to improve the product in ways that aren't literally training, such as quality reviews carried out by people.
The third is where the ambiguity lives. "We don't train on your data" is compatible with an engineer reading your support conversations to debug a ranking problem. That may be entirely acceptable to you. It is not acceptable to discover it later. Ask for it in the contract, not the sales deck.
The answer that should worry you: "we don't train on customer data by default." Those two words mean there is a setting, and you need to know who can change it. Also worrying: an answer about the vendor's own models that goes quiet on everyone else's.
The answer that should reassure you: a written commitment covering the vendor and every AI provider in the path, and a clear statement of what human access exists for support and debugging, with the controls around it.
10. Whose systems your questions pass through
This is the subprocessor question, and it is sharper for AI vendors than for ordinary software. You are not only asking who stores your material. You are asking whose systems your people's questions and your content pass through at the moment an answer is produced.
Get the actual list: which providers, for which parts of the job, hosted where, under what terms. Many products use different providers for different tasks, and each is a separate path your content travels. Then ask the question your security team will ask if you don't: how do you find out when it changes? Ask about region and hosting too, and treat vague answers as no. Residency is only worth what is written into the agreement you sign; vendors without it will talk about their cloud provider's global footprint instead, which is a marketing fact rather than a promise to you.
The answer that should worry you: "we use industry-leading models." "We're model-agnostic." Both are product statements dressed as security answers. Also worrying: a subprocessor list naming cloud providers but no AI providers.
The answer that should reassure you: a published, versioned list naming them, notification terms with a defined window and a right to object, and a straight answer about what is kept versus what merely passes through.
11. What the record keeps, and whether you can export it
You will want this on the day something goes wrong, and on that day you will want it in your own tooling rather than in somebody's admin console. Three sub-questions: what is recorded, only sign-ins and admin actions or every question asked and the material each answer drew on; how long it is kept and whether you can change that; and whether you can get it out automatically or only as a spreadsheet someone generates on request. Without a record of what an answer drew on, you cannot reconstruct a disclosure. You will know a question was asked and answered, and you will not know the one fact the conversation is about.
Ask for the past tense, too. Reviewers rarely ask "can this person see this document," because they can check that themselves in thirty seconds. They ask who could see it in March, what has changed since, and who changed it.
The answer that should worry you: "audit logs are available in the admin panel." "We log all activity," with nothing about retention. Any product where the history is overwritten rather than kept.
The answer that should reassure you: a kept history showing what moved from what to what, by whom and when, with changes to access sitting on the same timeline as changes to content, and an export you can automate.
12. What happens to your material when you leave
Your legal team will ask about retention and deletion regardless, so spend your question on the part they might not reach: everything derived from your content rather than the content itself. Deleting your documents is straightforward. Deleting the working copies, the stored answers, the logs containing quoted passages, the sample sets and the backups is not. "We delete customer data within 30 days of termination" often means the original documents only. Ask about deletion during the relationship as well: when an article is deleted in your own system, how quickly does the assistant stop using it? If it can still quote a document that no longer exists, that is a real problem for anything legally sensitive.
The answer that should worry you: a commitment naming only "customer data" without defining it, no mention of derived material, logs or backups, and no stated timeline.
The answer that should reassure you: a defined scope covering everything derived from your content, a timeline including backup expiry, an export path for your content and configuration, and a stated lag between deleting something and the assistant ceasing to use it.
13. What it does when it doesn't know
This lands last because it isn't strictly a security question. It is here because it is the failure your users will actually experience, and because it tells you something about the vendor that the other twelve don't.
Every assistant will sometimes find nothing relevant. What happens next is a choice: say it doesn't know and route to a person, or answer from general knowledge, producing something fluent and plausible that isn't grounded in anything you wrote. From the outside those two outputs are hard to tell apart, which is the entire problem. Ask whether answers are held to your own material, whether every claim carries a citation the reader can open, and how sure it has to be before it declines. Then ask whether you can tune that. A confidently wrong answer about your refund policy is a ticket, an escalation and sometimes a commitment you have to honour. We've gone deeper on the pattern in why AI gives confident wrong answers.
The answer that should worry you: "the model is very accurate." "It works from your content, so it can't make things up." Working from your content reduces invention. It does not eliminate it, and any vendor claiming otherwise is either overselling or hasn't looked.
The answer that should reassure you: a described fallback, citations on every answer, a threshold you can adjust, and a straight admission that it will sometimes get things wrong.
Three tests to watch, not hear about
You are not being asked to evaluate anyone's architecture, and you should refuse to be drawn into it. You are being asked to tell a precise answer from an evasive one, a skill you already have from every other vendor conversation you have run. And one part of this you can check for yourself, in about ninety seconds, without knowing anything technical at all. Insist on watching rather than being told. Ask them to do three things in front of you, live, on a running system rather than a slide:
- Use the assistant as someone who has been given no access at all.
- Use it as someone whose access was switched off a moment ago.
- Use it as a member of the public, with no sign-in.
Three empty results is a pass. Anything else is a finding you got for free, before you signed anything. A vendor who would rather show you the roadmap than run the test has answered a different question, and you should note which question they preferred.
Then run the harder version yourself during the pilot, because good answers earn a pilot rather than trust. Restrict a document to one group, sign in as somebody outside it, and try to make the assistant reveal it: its title, its author, a paraphrase of what is in it, a count of how many documents like it exist. Repeat with no sign-in. Do that before you sign, and again after every significant release.
The decisions that are yours, not the vendor's
Good answers to the questions above describe what a product can do. These are the choices you will be making regardless of who you buy from, and making them early is most of what separates a six-week review from a two-week one.
Group people by why they get access, not by the org chart
Mirroring the reporting structure feels natural and ages badly. Teams reorganise every few quarters, access never followed reporting lines anyway, and contractors, partners and the one solutions engineer who needs pricing data break the pattern in the first week.
Build groups that describe a reason: Tier 2 Support, Reseller Partners, Beta Customers, Revenue Operations. The test is whether you can finish the sentence "members of this group may see X because Y." If the sentence needs an "and also", you want two groups. Ten to thirty well-named groups covers most companies your size. Two hundred means someone is modelling individuals, and you will be maintaining it personally. This is also a business decision teams habitually defer to IT, and it shouldn't be. You know which content is customer-safe. Your identity team knows how to represent those groups. Neither of you can do the other's half.
Who someone is and who content is for are different questions
A group describes a person. An audience describes a piece of content. They are easy to conflate, and conflating them turns every publishing decision into a permissions ticket routed through you.
Keep them separate. Your enablement lead publishing a troubleshooting guide should be able to say "this is for Partners and Tier 2" without changing anyone's group membership. Keep the mapping small enough to read on one screen. If explaining who sees what needs a spreadsheet, the security review is going to take a long time, and you will be the one in it.
Decide visibility per record and per field, not just per record
Some decisions are about whole categories: whether a role can touch escalations, contracts or incident write-ups at all. Then there is the one everyone pictures, whether this specific article or ticket is visible to these groups.
The one teams skip and later regret is the level below that: same record, different parts. A partner sees the case status and the resolution notes but not the internal escalation thread, the cost-to-serve figure or the customer's direct phone number. Without it, your team's workaround is to keep an internal copy and an external copy of the same record, and those copies drift apart, and the drift becomes its own incident. Most teams need all three levels eventually; the mistake is starting at the finest one everywhere. Start coarse, refine where the content is genuinely sensitive, and confirm the fine control exists before you need it, because the alternative is a duplication habit you will be unpicking for years.
Decide what "public" means, deliberately
Anything a visitor with no sign-in can reach should be public because a person chose to publish it, not because nobody got round to restricting it. That is the difference between a system where a mistake makes content invisible and one where a mistake makes it visible. Pick the first kind.
Then write down what each kind of visitor gets, in a sentence each. A member of the public sees only what was deliberately published, and nothing else is reachable, including by a lucky phrasing. A signed-in customer sees that plus what their account entitles them to, plus their own records. A partner adds their partner content and usually loses sight of commercial figures. An internal agent gets their team's content, and internal-only is simply another audience.
What to walk into the review holding
Diagrams do not pass reviews, and neither does a vendor's architecture deck. Evidence does, and you can gather all of it without writing a line of anything.
Bring the side-by-side table. Ask the same five real questions as each of your four audiences and put the answers next to each other. Five questions, four different sets of results. That table is the single most persuasive thing you can hand a reviewer, because it shows the boundary instead of describing it. It is also a fair thing to ask a vendor to produce, and a revealing thing to watch them attempt.
Bring the three failure tests, run live in front of you, all returning nothing, and the plain-English statement of which of your existing groups sees what, on one screen.
And bring the history, which is the one that gets underestimated. A record that is kept rather than overwritten, showing what moved from what to what, by whom and when. Changes to access on the same timeline as changes to content, because a document becoming visible and a person joining a group are the same event as far as risk is concerned. And an export, because the reviewer wants the data in their own tooling, not yours. Agree with your security team in advance where those records land, how long they are kept and who reads them, because an exportable history is only useful if somebody exports it.
Produce those four and the conversation is short. Skip them and no amount of architecture description will substitute, and the calendar will make the decision for you.
Where MatrixFlows fits
Plenty of search and knowledge tools decide what to hide only after they have already found it, because the product was built before anyone thought to ask. The questions above will tell you which kind you are looking at, and you should point them at us exactly as you would at anyone else.
MatrixFlows settles what a person is allowed to see before it goes looking, so material outside that is never retrieved, never summarised and never quoted back. Anything it cannot positively confirm someone is entitled to returns nothing, and a visitor who isn't signed in sees only what was deliberately published to the public. Access follows the sign-in system you already run, so removing someone once removes them here too, and you can set visibility by category, by record and by individual fields within a record. The history of who could see what, and who changed it, is kept and can be exported for your reviewer. Our notes on grounding AI in company knowledge and guardrails for customer-facing AI cover the rest of the setup.
What stays hard, and what you should say about it
Settling access before anything is retrieved removes a category of failure. It does not remove the work, and a reviewer who has done this before will go straight to the parts you glossed over. Saying these things yourself is worth more than being caught not knowing them.
Your existing sharing mistakes come along for the ride
If the assistant draws on Confluence, Drive, your ticketing system or your CRM, it inherits their access rules along with their content. A folder mistakenly shared with "everyone at the company" in 2022 is mistakenly shared in your assistant too, faithfully and instantly. It is executing someone else's error correctly.
Timing compounds it, and this is not a small effect. Microsoft's own documentation for Restricted Content Discovery is refreshingly direct: for SharePoint sites with more than 500,000 items, a change to what can be discovered "could take more than a week to fully process". The same page notes that the control "doesn't change existing permissions" and "doesn't remove content from the search index". Being harder to find and being off limits are different things, and vendors will sometimes offer you the first while you are asking for the second.
Then there is the accumulated debris. Those 88% of organisations carrying stale but still-enabled accounts have accounts that still belong to live groups. They pass every check, because as far as any system is concerned they are legitimate people. No product design fixes that. Cleaning up your access does.
An answer can reveal what no single document does
This is the boundary nobody holds completely, and you should present it as such rather than let a reviewer find it. Ten individually permitted facts can support a conclusion nobody intended to permit. "Which accounts is the CS team spending the most time on" can be answered from ordinary ticket volumes, and the answer is a churn list. No single item is sensitive. The picture is.
What you can say is that the assistant only ever works from what that specific person was entitled to see; that answers prepared for one person are not reused for another; that unusual, sweeping questions are worth watching in your logs; and that what someone concludes from material they were legitimately allowed to read is a governance question rather than a product one. The honest line is that you constrain what goes in, not what someone infers from it. Reviewers respect that sentence. They do not respect finding out you avoided it.
It will enforce your access data exactly, including where it is wrong
If your group membership is wrong, the assistant will apply the wrong rule quickly, consistently and at scale, which is arguably worse than applying it sloppily. Getting the product choice right moves the failure from "the tool leaked" to "our records were wrong", and only one of those can be fixed with a calendar reminder.
So put group membership on a review cycle, quarterly at minimum, and treat it as a security control rather than an admin chore. Treat a change to who can see something as seriously as a change to the thing itself. Keep the number of groups small enough that a person can read the list. And use the cheapest control there is, which is leaving out the sources nobody should be asking about in the first place. A system your assistant never draws on cannot leak through any of the five routes.
These questions test a description, not a system
A vendor can answer every question perfectly and still have something wrong in the one place you didn't look. That is what the three tests are for. The questions also tell you nothing about the vendor's own security posture, which is what the audit report is genuinely for, and nothing about obligations specific to you: if you are regulated, if you handle health or payment data, if you have promised your own customers where their data sits, none of that is on this list.
So: you bring the questions and get real answers to all of them, in writing, from someone technical on the vendor's side. Your security team threat-models the deployment against your environment, reads the contract language behind the verbal commitments, and tests a live instance adversarially. You have done the part that requires knowing the product. They do the part that requires knowing your risk.
Ask the questions before the review does
If the argument holds, the implication is uncomfortable but useful. An assistant that finds first and tidies up afterwards will eventually show someone something it shouldn't, and no policy you write will prevent it, because the problem is in the shape of the product rather than in anyone's behaviour.
You also will not find out on a schedule. This kind of failure does not page anyone. It surfaces months later, when someone forwards an answer to a colleague who wasn't supposed to receive it, and by then the question is not whether it happened but how long it had been happening, and what you knew when you signed.
None of that requires you to become technical. It requires you to ask thirteen questions, watch three tests, and bring four pieces of evidence to a meeting. You can do the whole thing in an afternoon, and you can do it before the review rather than during it. Set up a workspace, connect a couple of real sources, create two groups with genuinely different access, and ask the same question as each. Then ask it as someone who isn't signed in, and confirm you get nothing back. That is a better vendor evaluation than any questionnaire, including this one.
Create a Free Workspace →
Related reading