Ātra minēšana bez avota
The answer with no handle
The meeting went quiet after the assistant produced the answer. That was the first warning. People usually make noise when software fails in an obvious way. They lean back, sigh, ask who owns the thing, and develop sudden opinions about procurement. This answer did not fail obviously. It was fluent, tidy and exactly the kind of paragraph that makes a project team sit up straighter. It said the policy allowed the exception. It gave three reasons. It used the same vocabulary as the policy office. It even sounded slightly bored, which is how institutional truth often dresses for work.
Then someone asked where the answer came from. The screen offered a source label that said internal guidance. That was not a source. That was a mood with a filing cabinet. Which guidance. Which version. Which paragraph. Was the document still active. Did the questioner have permission to see it. Had the text been copied from a draft, a current policy, a meeting note, or a training example produced by another system. The assistant could not say. It had retrieved something, or claimed it had. The team had an answer with no handle. It could be admired, but it could not be lifted.
This is the quiet trap in retrieval augmented AI. Retrieval makes a model feel grounded because the answer is no longer coming from the model alone. A search layer finds documents, chunks, records or facts, passes them into context, and the model writes from there. That is useful. It is also dangerously easy to overrate. If the system cannot show what it retrieved, why it was allowed to retrieve it, how fresh it was, which transformations touched it and how the final answer depends on it, retrieval has not solved guessing. It has put guessing on a faster route.
The old model guessed from memory. The new system may guess from a pile of fragments. That is an improvement only when the pile is governed. Otherwise the organisation has built a confident librarian who runs very quickly through the archive while refusing to keep shelf numbers. Impressive cardio. Poor audit trail.
Retrieval is a transport layer, not a truth layer
The simplest way to misunderstand retrieval is to treat search results as truth. Search does not know truth. Search knows match. Vector search knows similarity. Keyword search knows terms. Hybrid search knows a negotiated truce between the two. A reranker can improve ordering. Metadata filters can remove obvious mistakes. None of those steps automatically knows whether a paragraph is current, authorised, complete, contradicted, superseded, confidential or written by someone who was trying to get home before the train strike.
That does not make retrieval weak. It makes retrieval specific. It is a transport layer that brings candidate evidence to a reasoning surface. Its job is to find possibly relevant material under constraints. The truth work begins when the system can identify the material, preserve its context, compare it with other material, reject stale or unauthorised sources, expose uncertainty and keep a record of what happened. Without those pieces, retrieval is only faster selection. Faster selection is not the same as better judgement. A coin tossed by a machine is still a coin toss, even if it uses cloud credits.
In practical systems, the gap appears in small places. A source document is split into chunks, but the chunk loses the heading that made it meaningful. A vector index contains both current and retired policies because deletion was handled by enthusiasm rather than process. A PDF was parsed without footnotes. A table became a string of lonely numbers. An access control rule was applied to the document but not to the embedding. The answer cites a paragraph that was relevant to one region and illegal in another. Nobody designed a failure. They just allowed evidence to shed its labels as it moved.
The cure is not to distrust retrieval. The cure is to stop asking it to perform jobs it was never built to do. Retrieval should bring candidates. Provenance should tell the story of those candidates. Governance should decide which candidates may be used. The answer should carry enough evidence that people can inspect it without becoming amateur archaeologists in a folder called archive final old.
The missing shelf number
Libraries understood provenance before AI rediscovered it with more expensive terminology. A useful citation lets a reader find the work, edition, page and sometimes even the paragraph. It does not merely say history book. A warehouse label tells you which batch, supplier, lot and expiry date matter. A lab sample carries chain of custody because nobody wants medicine based on vibes. In each case, the point is mundane: when consequences exist, objects need identity across movement.
Data moving through retrieval systems needs the same discipline. A document is not a blob of text. It has an origin, owner, legal basis, scope, audience, version, effective period, format, parser, chunking policy, embedding model, index time, access policy and retirement path. That sounds like a lot because it is. It is still less than the cost of explaining to a regulator, patient, customer or board that the system probably read something helpful but the exact something has achieved spiritual independence.
The shelf number also protects engineering teams. When an answer is wrong, the team needs to know whether retrieval missed the right source, ranked the wrong source, included stale material, lost context during chunking, allowed a permission leak, or let the model ignore the best evidence. Those are different bugs. Without provenance, they collapse into one useless category: AI was weird. That category is popular in meetings and almost completely unhelpful in incident review.
Good provenance gives failure a shape. It lets teams inspect the document path, the chunk boundary, the embedding version, the query rewrite, the rerank result, the prompt assembly and the final generation. It does not make the system perfect. It makes mistakes locatable. Locatable mistakes can be repaired. Unlocatable mistakes become folklore, and folklore has terrible uptime.
Chunking is an editorial act
People often describe chunking as a technical preprocessing step. It is more than that. Chunking decides what context travels together. A paragraph split away from its exception note can change meaning. A warranty clause separated from its jurisdiction heading becomes a trap. A clinical recommendation without the patient group that limits it is no longer the same recommendation. A code sample without its warning is an ambush wearing monospace.
Every chunking strategy is an editorial policy. Fixed token windows are simple and fast, but they can cut through meaning. Semantic chunking respects topic shifts, but it can still miss tables, lists, captions and legal structure. Section-based chunking preserves hierarchy, but only if the parser actually recognises the hierarchy. Parent-child retrieval can bring a small match and a larger surrounding passage, but then the system must record both. None of this is mystical. It is document work. Document work is where many AI systems discover that the boring office had been protecting meaning all along.
Provenance turns chunking from silent damage into a visible choice. The system should know which document produced a chunk, which parser read it, which headings enclosed it, which sibling chunks were nearby, which page or row it came from, and what larger unit can be shown when a human challenges the answer. The answer does not need to display all of that every time. It does need to keep it. The difference between hidden context and absent context becomes important at precisely the moment everyone is already tired.
Chunking also affects fairness across sources. A clean HTML page may produce beautiful chunks. A scanned PDF may produce broken fragments. A spreadsheet may lose structure. If the retrieval system quietly favours the sources that parse neatly, it may also favour departments, suppliers or languages with better document hygiene. That is not model bias in the dramatic sense. It is office bias, which is less cinematic and surprisingly powerful.
Citations are not enough
A citation can be a useful user interface for provenance, but it is not the whole mechanism. Many systems show little source chips beside an answer. That is better than nothing. It is also easy to game by accident. The answer may cite a document that was retrieved but not actually used. It may cite a page that contains similar words but not the claim. It may cite a source that was available to the system but not to the user. It may cite an old version because the index lags behind the repository. A decorative citation is still decoration, only with a footnote costume.
Nopietna izcelsmes dokumentācija prasa disciplīnu apgalvojumu līmenī. Ja atbilde saka, ka politika pieļauj izņēmumu, sistēmai būtu jāzina, kura rindkopa pamato šo apgalvojumu. Ja tā saka, ka izņēmums attiecas tikai zem noteikta sliekšņa, tai būtu jānorāda sliekšņa avots. Ja tā apvieno divus avotus, tai būtu jāsaglabā savienojums redzams. Ja pierādījumi ir pretrunīgi, tai nevajadzētu saplacināt konfliktu vienā jautrā rindkopā. Modelis var apkopot, bet sistēmai nevajadzētu ļaut kopsavilkumam izdzēst pierādījumu struktūru.
Tas nenozīmē, ka katrai atbildei jānāk ar juridisku dokumentu komplektu, kas sasiets ar aukliņu. Dažādi konteksti prasa dažādu virsmas detalizācijas līmeni. Atbalsta dienesta atbilde var parādīt divas avotu saites un pārliecības piezīmi. Medicīnas, finanšu, juridiskā vai sabiedrisko pakalpojumu darbplūsma var prasīt rindkopu atsauces, versiju identifikatorus un cilvēka pārskatīšanas statusu. Galvenā prasība ir tāda, ka virsma var paplašināties, kad pieaug sekas. Sistēma, kas nevar pāriet no vienkāršas atbildes uz pārbaudāmiem pierādījumiem, ir tērzēšanas robots nopietnā kreklā.
Atsaucēm vajag arī negatīvo telpu. Sistēmai būtu jāspēj pateikt, ka tā neatrada pietiekami daudz pierādījumu, vai ka izgūtie avoti ir pretrunīgi, vai ka avoti ir novecojuši, vai ka lietotājam nav piekļuves nepieciešamajam materiālam. Atteikums ar izcelsmes dokumentāciju bieži ir noderīgāks par atbildi ar viltotu atsauci. Organizācijai var nepatikt dzirdēt nē. Organizācijām reti patīk. Tieši tāpēc pastāv pārvaldība, un reizēm kafija.
Aktualitāte ir daļa no patiesības
Izcelsmes dokumentācija bez laika ir nepilnīga. Daudzas uzņēmumu kļūdas rodas no veca materiāla, kas paliek meklējams, jo neviens negribēja dzēst neko ar virsrakstu, kas izklausījās svarīgs. Politikas beidzas. Cenas mainās. Produktu rokasgrāmatas tiek aizstātas. Regulatīvās vadlīnijas mainās. Datu vārdnīcas noveco. Izgūšanas sistēma, kas vecu un aktuālu materiālu traktē vienādi, nav neitrāla. Tā nodod laika pārvaldību kosinusa līdzībai, kas ir drosmīga dzīvesveida izvēle.
Katram avotam izgūšanas sistēmā būtu jānes laika nozīme. Kad tas tika izveidots. Kad tas stājās spēkā. Kad tas pēdējo reizi tika pārskatīts. Kad tas tika indeksēts. Kad tas beidzas. Kura versija to aizstāja. Vai tas bija melnraksts, apstiprinājuma kopija, arhīva kopija vai aktuālā kopija. Tās nav birokrātiskas rotas. Tās ir daļa no tā, vai atbilde ir pietiekami patiesa, lai to izmantotu. Rindkopa no pagājušā gada procedūras var būt perfekti uzrakstīta un pilnīgi nepareiza.
Aktualitātei ir arī operacionālas sekas. Indeksēšanas kavējumiem jābūt redzamiem. Ja repozitorijs mainījās pulksten 09:00 un vektoru indekss atjaunojas nakts laikā, sistēmai jāzina par šo plaisu. Ja steidzams materiāls apiet parasto cauruļvadu, apiešana jāreģistrē. Ja avots tiek izbeigts, iegultnes un atvasinātie fragmenti jāizbeidz kopā ar to vai arī tie jāmarķē kā vēsturiski. Pretējā gadījumā sistēma kļūst par muzeju, kas reizēm sniedz operacionālus padomus.
Ar laiku saistīta izcelsmes dokumentācija palīdz lietotājiem uzticēties pareizajām lietām. Tā ļauj asistentam pateikt, ka šī atbilde ir balstīta uz politiku, kas ir spēkā no 2026. gada 5. marta, indeksēta pulksten 11:20, un nav atrasts jaunāks aizstājošs ieraksts. Šis teikums nav glamūrīgs. Tas ir noderīgs. Noderīgs pārspēj glamūru katrā incidentā, ko esmu saticis.
Atļaujas ceļo arī
Piekļuves kontrole bieži tiek piemērota pie priekšējām durvīm un aizmirsta gaitenī. Lietotājam var nebūt atļauts atvērt avota dokumentu, bet dokumenta iegultne var atrasties koplietotā indeksā. Fragments var tikt iekļauts uzvednē, jo izgūšanas pakalpojums darbojas ar plašu pakalpojumu kontu. Ģenerēta atbilde var atklāt konfidenciālas lietas esamību, pat to necitējot. Sistēma neizpludināja failu, kāds saka. Tā tikai izpludināja secinājumu. Šī atšķirība vislabāk tiek pasniegta no droša attāluma.
Izcelsmes dati ir jāiekļauj atļaujas, jo pierādījumi bez pilnvarojuma nav izmantojami pierādījumi. Sistēmai būtu jāzina, kurš lietotājs, kāda loma, kāds mērķis un kāds konteksts ļāva katram iegūtajam vienumam nonākt atbildē. Tai būtu jānošķir piekļuve avotam no atvasinātas piekļuves. Tai vajadzētu veikt rediģēšanu pirms ģenerēšanas, ja nepieciešams. Tai būtu jāreģistrē, kad atbilde bija ierobežota atļaujas dēļ, nevis pierādījumu trūkuma dēļ. Pretējā gadījumā lietotāji klusumu var pārprast kā neesamību, vai vēl ļaunāk, saņemt informāciju, kuru viņiem nekad nebija paredzēts redzēt.
Ar atļaujām saskaņota izguve ir grūtāka nekā parastā izguve, jo tā maina ranžēšanu, kešošanu, novērtēšanu un testēšanu. Divi lietotāji var uzdot vienu un to pašu jautājumu un pamatoti saņemt atšķirīgus pierādījumus. Tā nav nekonsekvence. Tā ir pārvaldība. Grūtākā daļa ir padarīt atšķirību izskaidrojamu, neatklājot to, kam jāpaliek slēptam. Sistēmai var nākties pateikt, ka ārpus jūsu piekļuves loka var būt ierobežoti ieraksti, nevis izlikties, ka pasaulē pastāv tikai tas, ko lietotājs var izlasīt.
Šeit arī daudzi pilotprojekti izgāžas, saskaroties ar realitāti. Prototips, kas balstīts uz koplietojamu mapi, var nedēļu visus pārsteigt. Tad kāds jautā par personāla lietu ierakstiem, iegādes dokumentiem, pacientu piezīmēm, juridisko privilēģiju, darbinieku pārstāvības institūcijas materiāliem vai eksporta kontrolētu pētniecību. Izguves sistēmai pēkšņi nepieciešama pieaugušo uzraudzība. Jautrā demonstrācija kļūst par datu pārvaldības projektu, kas tā arī bija visu laiku.
Pretruna nav kļūda, ko slēpt
Reāli arhīvi ir pretrunīgi. Politikas komanda atjaunināja procedūru, bet ne bieži uzdoto jautājumu sadaļu. Bieži uzdoto jautājumu sadaļa atjaunināja piemēru, bet ne tabulu. Reģionālais birojs saglabāja vietējo izņēmumu. Līgumā teikts viens, ieviešanas rokasgrāmatā cits, un izklājlapa, ko izveidojis ļoti praktisks cilvēks, parāda, ko visi patiesībā dara. Izguve atradīs visu, ja vaicājums būs neveiksmīgs vai godīgs.
Sistēmai, kas apzinās izcelsmi, būtu pretrauna jāuztver kā pirmšķirīgs rezultāts. Tai būtu jāparāda, ka vairāki avoti ir pretrunā, jānorāda to pilnvarojums un aktualitāte, un jāizvairās pasniegt vienu sapludinātu atbildi tā, it kā organizācija būtu runājusi vienā balsī. Dažreiz pareizā atbilde nav tāda, ka izņēmums ir atļauts. Dažreiz tā ir tāda, ka pašreizējā politika, šķiet, to pieļauj, bieži uzdoto jautājumu sadaļa, šķiet, ir novecojusi, un līguma īpašniekam būtu jāatrisina konflikts pirms rīcības. Šī atbilde ir mazāk ērta. Tā arī mazāk radīs nelielu juridisku laikapstākļu sistēmu.
Pretrunīgu informāciju apstrādājot, avoti jākārto pēc autoritātes, nevis tikai pēc atbilstības. Valdes apstiprināta politika var būt svarīgāka par palīdzības rakstu. Parakstīts līgums var būt svarīgāks par pārdošanas prezentāciju. Vietējā procedūra var būt svarīgāka par vispārīgo rokasgrāmatu savā vietējā darbības jomā. Projekts nedrīkst būt svarīgāks par spēkā esošu versiju, ja vien lietotājs tieši nejautā par projektiem. Šie noteikumi ir garlaicīgi. Tie ir arī vieta, kur organizācijas patiesības hierarhija kļūst tehniska.
Ja neviens nevēlas definēt šo hierarhiju, izguves sistēma to definēs nejauši. Tā izmantos teksta līdzību, aktualitāti, formatējuma kvalitāti, fragmenta garumu vai jebkuru citu signālu, ko piedāvā cauruļvads. Nejauša autoritāte joprojām ir autoritāte. Tā vienkārši nonāk bez protokola.
Novērtējumā jāiekļauj avota kļūmes, ne tikai atbildes kļūmes
Daudzas komandas izguves sistēmas novērtē, pārbaudot, vai galīgā atbilde izklausās pareizi. Tas ir noderīgi, bet nepietiekami. Pareiza atbilde no nepareiza avota ir nākotnes incidents, kas tikai sasilst. Sistēma var būt guvusi panākumus tikai tāpēc, ka modelis jau zināja atbildi, vai tāpēc, ka novecojis dokuments nejauši sakrita ar pašreizējo noteikumu, vai tāpēc, ka vērtētājs pieņēma atsauci, kas neatbalstīja apgalvojumu. Atbildes kvalitāte un pierādījumu kvalitāte jāpārbauda atsevišķi.
Izguves novērtējuma komplektā jāiekļauj avotu sagaidāmie rezultāti. Katram testa jautājumam, kuri dokumenti ir pieņemami. Kuri nav pieņemami. Kurām versijām ir nozīme. Kuras atļaujas attiecas. Kuri konflikti jāizceļ. Kura atbilde jāatsaka, jo trūkst pierādījumu. To izveidot ir lēnāk nekā jautājumu un atbilžu pāru kaudzi. Tas ir arī tuvāk darbam, kas sistēmai jāveic. Izguves sistēma bez avotu novērtējuma ir kā finanšu sistēma, ko pārbauda tikai jautājot, vai galīgais skaitlis izskatās ticams. Tā var izturēt pārbaudi līdz pat auditam.
Operatīvajā uzraudzībā jāvēro arī avotu uzvedība. Kuri avoti ir pārstāvēti pārāk bieži. Kuri avoti tiek izgūti reti, bet bieži ir nepieciešami. Kuri fragmenti tiek citēti bieži. Kuras atbildes vēlāk tiek labotas. Kuri novecojuši dokumenti turpina parādīties. Kuras lietotāju grupas saņem vairāk atteikumu, jo atļauju robežas ir slikti modelētas. Šie signāli nav tikai tehniski rādītāji. Tie ir pierādījumi par zināšanu krājuma veselību.
Kad izcelsme ir pieejama, novērtējums kļūst precīzāks. Neveiksmīgu atbildi var izsekot līdz izguvei, ranžēšanai, avota kvalitātei, piekļuves politikai, uzvednes veidošanai vai ģenerēšanai. Katrai kļūmju klasei ir cits īpašnieks. Tas ir neērti, jo neļauj mierinošajai frāzei mākslīgais intelekts to sajauca norīt visu problēmu. Labi. Komforts ir pārvērtēts, ja sistēma pieņem lēmumus.
Lēmumu cikls
Izcelsme nedrīkst būt arhīva funkcija, ko pievieno beigās. Tai jābūt lēmumu ciklā. Lietotājs jautā. Sistēma izgūst saskaņā ar politiku. Atbilde nes pierādījumus. Lietotājs pieņem, apstrīd vai labo. Labojums atjaunina avota kvalitāti, ranžēšanas noteikumus, metadatus, piekļuves kontroli vai apmācības piemērus. Nākamā atbilde netiek vienkārši ģenerēta no jauna. To nosaka tas, ko organizācija ir iemācījusies.
Šis cikls ir tas, kā izguve kļūst par institucionālo atmiņu, nevis gudru dokumentu automātiskās pabeigšanas rīku. Bez cikla katra atbilde ir atsevišķs notikums. Ar ciklu atbildes kļūst par signāliem par zināšanu sistēmas stāvokli. Slikta atbilde var atklāt novecojušu politiku. Atteikums var atklāt trūkstošu dokumentāciju. Pretruna var atklāt neatrisinātu īpašumtiesību jautājumu. Biežs jautājums var atklāt, ka procedūra ir nesalasāma. Izguves slānis kļūst par diagnostikas instrumentu, ne tikai atbilžu mašīnu.
The loop also gives humans a sane role. People should not be asked to inspect every token. They should be asked to resolve the meaningful failures that provenance exposes. Is this source authoritative. Is this exception current. Is this access boundary correct. Is this conflict real. Those are human governance questions. The system can route them, record them and learn from their answers. It should not bury them under fluent text.
In the end, provenance is not an academic nicety. It is the difference between an AI system that can participate in accountable work and one that can only sound helpful until challenged. Retrieval gets material into the room. Provenance says who brought it, from where, under what authority, and whether anyone should trust it enough to act.
The lesson
Retrieval is one of the most useful patterns in applied AI because it connects models to living knowledge. That usefulness is exactly why it needs provenance. The more people rely on retrieved answers, the less acceptable it becomes to say the source was somewhere in the index. Somewhere is not a control. Somewhere is where bad meetings begin.
A serious retrieval system keeps identity across movement. It preserves source, version, permission, time, chunk context, authority, contradiction and use. It evaluates whether evidence supports claims, not only whether answers sound good. It makes refusal possible when proof is thin. It lets people challenge and repair the knowledge estate. This is not paperwork around AI. It is the part that turns retrieval from fast guessing into accountable assistance.
The model may write the answer. The retrieval layer may find the words. Provenance is what lets the organisation own the claim. Without it, all that speed only gets uncertainty to the user sooner.