Sabiedriskais internets kļūst par apmācības datu sarunu platformu

Crawleri, rezervācijas, licences un izcelsmes apliecinājumi maina veco pieņēmumu, ka publiska lapa vienkārši ir paredzēta pārņemšanai. Noderīgais jautājums...

Sabiedriskais internets kļūst par apmācības datu sarunu platformu

The page is public. The use is not settled.

For a long time, the public web ran on a blunt but workable social bargain. A publisher made a page available. Search engines found it, indexed enough of it to send readers back, and a person arrived at the original page. There were arguments about indexing, snippets and advertising, but the route was recognisable. A page was both the unit of publication and the place where the reader met the publisher.

Training a general-purpose model changes the shape of that bargain. A crawler may still begin with a public address, but the destination is no longer necessarily a search result. The material may be copied, parsed, filtered, stored, combined with other material and used in a training process whose output is not a link back to the page. It may later help answer a question without a reader visiting the original work. The fact that a page is visible is relevant. It is not the whole answer.

This is why the public internet is becoming a training-data negotiation. Not a single negotiation held in one room, and not a drama in which one side discovers that the other exists. It is a distributed negotiation conducted through law, licences, technical signals, contracts, standards, records and, where necessary, people speaking to each other. A rightsholder can reserve a mining use. A model provider can identify a reservation and decide not to collect. The parties can agree a licence. An organisation can document the path well enough that a later question is answerable. Or they can leave the path vague and discover that vagueness is expensive once the material has travelled.

None of this makes every public page unavailable for every computational use. The European copyright framework contains text and data mining exceptions, and the conditions matter. Neither does it make a robots file into a universal copyright instrument, or a licence into a simple tick box. The useful change is more modest: training data has to be treated as a relationship with terms, history and boundaries. That is a better description of what it already was. The new part is that the relationship is becoming harder to ignore.

The European Commission's work around the AI Act gives the change an administrative shape. Providers of general-purpose AI models have copyright-policy and training-content-summary obligations. The GPAI Code of Practice presents a voluntary route for providers that choose to follow it, including commitments concerning rights reservations and robots.txt. The Commission has also consulted on protocols for expressing reservations. These are not substitutes for copyright law, nor a declaration that every technical signal has the same legal effect. They are evidence that the journey from a rightsholder's instruction to a model provider's collection system is now a policy problem in its own right.

The story does not need an invented publisher, a dramatic crawler or a fictitious court hearing. It is visible in the ordinary points where a source changes hands. Someone decides whether a URL belongs in a collection. Someone decides what a reservation means. Someone decides whether a licence is sufficiently specific for training, retrieval, evaluation or redistribution. Someone decides which version of a policy applies when a source changes. The decisions can be quiet. They are still decisions.

A crawler is no longer only a visitor

Browsers, search crawlers, archival tools, accessibility services, academic researchers, monitoring systems and model-training pipelines can all request the same page. From the publisher's server, they may initially look similar: a request arrives, a response is sent, a log line is written. But their purposes, copying patterns and downstream consequences are different. Treating every automated request as identical is convenient for infrastructure and poor for governance.

A normal reader opens one article and may return another day. A search crawler creates an index intended to direct a reader back. A training pipeline can make a copy that is retained, transformed and aggregated with many other works before any model exists. A retrieval system may create a more direct commercial substitute for a visit. An evaluation set may preserve excerpts to test a system. The labels are not moral rankings. They are descriptions of uses that require different questions.

The first question is usually access. Did the collector reach the material through a lawful route? The second is the legal basis for the acts the system wants to perform. The third is whether the rightsholder has expressed a reservation where the applicable rule allows one. The fourth is whether a separate licence or contract supplies permission, conditions or limits. A fifth question, which often arrives too late, is operational: can the organisation show what it saw and why it proceeded?

Directive (EU) 2019/790 does not treat text and data mining as a single, limitless permission. Article 3 provides a mandatory exception for research organisations and cultural heritage institutions making reproductions and extractions for scientific research, provided they have lawful access. Article 4 addresses text and data mining for other purposes, again where there is lawful access, unless the rightsholder has reserved the relevant rights in an appropriate manner. For online content made publicly available, the Directive says the reservation may be expressed by machine-readable means, including metadata and terms and conditions.

That distinction is not a minor drafting curiosity. It means that a public page does not carry one generic label called usable. The actor, purpose, access route, rightsholder's reservation and the particular acts all matter. A sensible collection system therefore needs more vocabulary than allow and block. It should be able to record candidate, access observed, reservation detected, licence required, licence granted, purpose limited, unresolved, declined and withdrawn. This is not bureaucracy for its own sake. It is the minimum grammar needed when a source can be perfectly visible and still not be available for the proposed use.

The European Union Intellectual Property Office's 2025 study on generative AI and copyright describes this landscape as one in which rights reservations and licensing are developing alongside technical solutions. Its conclusion is not that a finished market mechanism has appeared. It is nearly the opposite: no single solution fits every rightsholder, every model provider or every kind of material. That is precisely why an honest system has to preserve distinctions rather than erase them in the name of scale.

Scale is often offered as an excuse for imprecision. A model provider may say it cannot negotiate page by page. A small publisher may say it cannot inspect every automated request. Both statements can be true. They do not remove the need for a workable interface between the two. They explain why machine-readable signals, standard licences, source-family records and collective arrangements are interesting. The answer to a large problem is not necessarily a large meeting. It may be a dependable way to state a boundary and a dependable way to receive it.

Avota maršruts nav tunelis uz modeli. Tā ir virkne lēmumu, kam būtu jāpaliek redzamiem arī pēc vākšanas.

Diagramma ir konceptuāls maršruts, nevis juridisks lēmumu dzinējs. Zaļgana kartīte nenozīmē, ka izmantošana ir likumīga. Tā parāda pierādījumus un spriedumu, ko komandai būtu jāspēj pārbaudīt. Šī pieticīgā atšķirība ir vērtīga. Lielākā daļa novēršamo pārpratumu sākas tad, kad operatīvs fakts tiek sajaukts ar juridisku secinājumu vai kad juridisku secinājumu nevar izsekot līdz operatīvajiem faktiem.

Atsauces prasa valodu, ko mašīnas spēj pārnēsāt

Opt-out bieži apraksta kā pogu: tiesību īpašnieks saka nē, pārlūkprogramma apstājas. Faktiskais darbs ir mazāk kinematogrāfisks. Atsauce ir jāizsaka kaut kur, kolektoram tā jāatrod, jāsaista ar attiecīgo darbu vai pakalpojumu, jāinterpretē pareizajā kontekstā un jāsaglabā kā pierādījums. Tā var būt metadatos, lietošanas noteikumos, publicētā politikā vai protokolā, ko kolektors saprot. Tā var attiekties uz tīmekļa vietni, katalogu, plūsmu, datubāzi vai noteiktu darbu grupu. Tā laika gaitā var mainīties.

Direktīvas atsauce uz mašīnlasāmu nozīmi ir svarīga, jo atsauce, ko nevar atrast vākšanas brīdī, nevar ticami mainīt vākšanas lēmumu. Bet mašīnlasāms nenozīmē nekļūdīgs, pašsaprotams vai juridiski pilnīgs. Parsētājs var atrast direktīvu ar nepazīstamu vērtību. Tiesību īpašnieks var publicēt politiku, kas norāda uz licencēšanas kontaktpersonu, nevis kategorisku atteikumu. Avots var būt pieejams ar abonementu, kam ir savi noteikumi. Sistēma var neredzēt nekādu atpazīstamu signālu. Signāla neesamība ir novērojums, ne visaptveroša atļauja.

Šeit ir diezgan nīderlandiešu mācība. Zīme, ka tilts ir slēgts, nekļūst labāka tāpēc, ka piegādes šoferis apgalvo, ka viņa karte zīmi neparsēja. Arī ne katrs slēgts tilts ir aizliegums iebraukt provincē. Jautājums ir, vai zīme ir salasāma attiecīgajam maršrutam, vai maršrutam ir saprātīga alternatīva un vai kāds ir pierakstījis, kas notika. Mērķis nav pārvērst tīmekli par kanālu tīklu. Mērķis ir neļaut robežai pazust tāpēc, ka to bija neērti modelēt.

Komisijas apspriešanā par protokoliem tiesību rezervēšanai tekstizraces un datizraces vajadzībām šī īstenošanas plaisa kļūst skaidri redzama. Tajā atsauce izdarīta uz GPAI prakses kodeksa apņemšanos, ka parakstītāji identificē un ievēro atbilstošus mašīnlasāmus protokolus, ko izmanto tiesību subjekti, kā arī uz apņemšanos ievērot robots.txt un turpmākās IETF standarta versijas. Pati apspriešana nav galīgais protokols. Tā ir pierādījums tam, ka pusēm ir nepieciešama skaidrāka tehniskā valoda juridiskai izvēlei, kurai jāceļo caur automatizētām sistēmām.

Robots.txt šajā sarunā ir noderīga, bet ierobežota loma. Tā ir sen izveidota tīmekļa konvencija pārlūkošanas programmu darbībai. Tā var norādīt, ka izdevējs nevēlas, lai konkrēts automatizēts aģents vai ceļš tiktu pārlūkots. GPAI kodeksa attieksme tam piešķir praktisku nozīmi parakstītāju vākšanas sistēmām. Taču robots.txt nav universāls tiesību reģistrs, un pārlūkošanas instrukcija neatrisina visus autortiesību, līguma vai datubāzes tiesību jautājumus. Šo līmeņu sajaukšana rada divus sliktus iznākumus. Viena puse pieņem, ka robots.txt ir pilnīga juridiskā atbilde. Otra to uzskata tikai par neobligātu etiķeti. Neviena no šīm pozīcijām nerada stabilas attiecības.

Vākšanas sistēmai būtu jātur signāli atsevišķi. Tehniskā pārlūkošanas instrukcija: ko robots fails norādīja attiecīgajā brīdī? Tiesību rezervēšana: kādu mašīnlasāmu vai publicētu paziņojumu sniedza avots? Piekļuves nosacījums: vai ceļš bija atvērts, autentificēts, abonēts vai nodrošināts saskaņā ar līgumu? Licence: kādas tiesības tika piešķirtas kādam mērķim un periodam? Iekšējā politika: ko organizācija nolēma darīt gadījumos, kad aina palika neskaidra? Šie lauki var būt saistīti, taču tie nav savstarpēji aizstājami.

Šī nošķiršana padara kļūdu apstrādi arī cilvēcīgāku. Vācējs var ziņot, ka redzējis signālu, ko nespēj interpretēt, nevis klusībā uzskatīt signālu par nebūtisku. Tiesību subjekts var atklāt, ka tā vēlamais marķieris netiek atbalstīts, un izvēlēties citu kanālu. Pakalpojumu sniedzējs var uzlabot parsētāju, nepārrakstot vēsturi. Vēlāks recenzents var redzēt, vai avots tika pieņemts tāpēc, ka pastāvēja licence, tāpēc, ka tika izvērtēts izņēmums, vai tāpēc, ka tika pieņemts politikas lēmums norādītās neskaidrības apstākļos. Alternatīva ir viens necaurredzams lauks, kurā rakstīts vākts.

Licencēšana nav inovācijas pretstats

Pastāv noturīgs ieradums licences raksturot kā berzi un modeļus kā progresu. Tas ir maldinošs pretstats. Licence var būt ierobežojums, bet tā var būt arī saskarne: veids, kā pateikt, kurš materiāls ir pieejams, kādam lietojumam, par kādu cenu, ar kādu attiecinājumu, atskaitēm, izslēgumiem vai atjaunošanas nosacījumu. Šīs saskarnes kvalitāte nosaka, vai mazāki tiesību subjekti un mazāki modeļu sniedzēji vispār var piedalīties.

Eiropas izdevēju federācija ir paudusi viedokli, ka tiesību subjektiem ir nepieciešama pārredzamība par datiem, ko izmanto mākslīgā intelekta apmācībai un no kurienes tie iegūti, vienlaikus atzīstot, ka izdevēji paši savā darbā var izmantot mākslīgo intelektu. Tā ir nozares pozīcija, nevis neitrāls pierādījums pareizajai juridiskajai atbildei katrā gadījumā. Tomēr tajā ir formulēts praktisks punkts, ko modeļu sniedzējiem būtu jāņem vērā: izdevējs var būt gan mākslīgā intelekta lietotājs, gan tiesību subjekts, kura materiālam ir nepieciešami noteikumi. Izvēle nav starp literatūru un tehnoloģiju vai starp radītājiem un inženieriju. Izvēle ir par to, vai apmaiņai ir saprotams pamats.

The Publishers Association's 2026 account of the UK publishing and AI licensing market provides a more specific view. It describes existing text and data mining licences, later AI-training licences and growing retrieval-augmented-generation licensing, and it reports work towards a collective licence with opt-in for rightsholders. The United Kingdom is outside the EU legal order, so this is not evidence of what the DSM Directive requires. It is useful evidence of a nearby European rights market attempting to make training and retrieval permissions more practical than a private negotiation between a handful of very large organisations.

Collective arrangements are worth watching because the web is not made only by companies with a legal department and an account manager. A small specialist publisher, a learned society, a regional newspaper, a photographer's archive or an independent author may have valuable material and very little capacity to negotiate bespoke terms with every potential model provider. A standard opt-in, a collective licence or a trusted intermediary cannot resolve every valuation question. It can make the first conversation possible.

Model providers also have a reason to prefer clarity. A negotiated route can supply material with an identified provenance, a defined scope and an accountable contact. It may cost money. So does cleaning a dataset, investigating a disputed source, defending a position that has no records, or replacing a source late in development. A licence does not guarantee that the material is suitable, accurate or representative. It makes a different contribution: it turns permission from an assumption into an expressed condition.

Compensation is not a magic word either. There is no single fair tariff for every work and every use. Training, retrieval, evaluation, internal analysis, model improvement and publication can create different commercial relationships. A licence might be per work, per collection, per volume, per period, per deployment, per user, per model family or negotiated on another basis. It may include an attribution duty, a reporting duty, a right to audit, a withdrawal mechanism or no continuing access at all. A useful system does not pretend these terms are universally simple. It makes them legible enough to apply.

The important shift is from a culture of extraction to a culture of terms. That does not mean every dataset becomes a procurement exercise. It means that when a rightsholder chooses to reserve, license or offer material under conditions, the model provider has an operational way to receive that choice. When a provider wants to use high-quality, current, specialised material, it has a route to ask rather than merely a route to copy. There is considerable room between indiscriminate crawling and a world in which only the largest firms can negotiate.

Provenance is the receipt, not the permission

Provenance is sometimes sold as a cure for copyright uncertainty. It is not. A source record can show where material came from and what somebody believed at a given time. It cannot make an unavailable work available. A hash can establish that two files are identical without explaining whether either was lawfully copied. A beautifully maintained ledger can document a bad decision with impressive precision.

That limitation is exactly why provenance is useful. It separates evidence from wishful thinking. If a model provider stores the source URL, the source identity, the collection time, the observed terms, the reservation evidence, the access route, the decision owner, the applicable licence or exception assessment, the dataset version and the downstream purpose, a reviewer can ask a meaningful question. If the provider stores only an extracted text fragment and a date, the reviewer is left with archaeology.

Laba izcelsmes uzskaite prasa laiku. Lapa var mainīties. Izdevējs var pārskatīt noteikumus. Avots var tikt noņemts, pārvietots aiz abonementa, pārdots, labots vai aizstāts. Licence var beigties. Atteikšanās protokols var tikt atjaunināts. Modeļa nodrošinātājs var atklāt defektu iepriekšējā kolekcijas lēmumā. Ieraksts nedrīkst pārrakstīt agrāko stāvokli ar pašreizējo. Tam jāsaglabā agrākais novērojums, jāidentificē vēlākās izmaiņas un jāreģistrē, ko organizācija darīja tālāk.

Tas ir īpaši svarīgi, ja tiek apspriesta noņemšana vai atsaukšana. Lapas noņemšana no tiešraides pārmeklēšanas rindas nav tas pats, kas tās noņemšana no katra starpposma krātuves, atvasinātās datu kopas, precizēšanas (fine-tuning) darbības, novērtēšanas kopas, izguves indeksa un izlaistā modeļa. Godīga atbilde nav solīt tūlītēju dzēšanu no katra tehniskā artefakta. Tā ir definēt ceļus, ierobežojumus un pārskatīšanas punktus, pirms tos solīt. Tiesību īpašniekam ir pelnījis kontaktpunktu, kas var izskaidrot procesu. Modeļa nodrošinātājam ir nepieciešams veids, kā identificēt, kuri artefakti ir ietekmēti. Abiem ir nepieciešams ieraksts, kas atšķir saņemtu pieprasījumu no atrisināta pieprasījuma.

AI akta 53. pants piešķir izcelsmes uzskaitei sabiedriskās politikas kaimiņu. Tas pieprasa vispārējas nozīmes mākslīgā intelekta modeļu nodrošinātājiem ieviest politiku, lai ievērotu Savienības autortiesību tiesību aktus, un publiskot pietiekami detalizētu kopsavilkumu par apmācībai izmantoto saturu. Publiskajam kopsavilkumam nav jāatklāj katrs datu kopas vienums, un akts ietver konfidencialitātes un komercnoslēpumu robežas. Bet virziens ir skaidrs: modeļa nodrošinātājam vajadzētu spēt aprakstīt apmācības saturu tādā veidā, kas ir informatīvāks par "uzticieties mums", vienlaikus saglabājot pilnīgāku dokumentāciju attiecīgajām iestādēm un pakārtotajiem nodrošinātājiem.

Šī auditoriju sadalījums ir saprātīgs. Publiskajam lasītājam ir jāsaprot avotu kategorijas, kolekcijas un kūrēšanas izvēles, attiecīgie ierobežojumi un kontaktceļi. Tiesību īpašniekam ar konkrētu jautājumu var būt nepieciešams process, kas spēj apstrādāt specifiskāku pieprasījumu. Regulatoram var būt nepieciešama dokumentācija, ko nevar saprātīgi publicēt publiskā lapā. Kļūda ir uzskatīt šos slāņus par attaisnojumu neko neteikt vai uzskatīt publisko kopsavilkumu par pierādījumu, ka katrs avota līmeņa lēmums ir atrisināts. Pārredzamība ir informācijas arhitektūra, nevis preses relīze.

Izcelsmes uzskaite nerada atļauju. Tā saglabā pierādījumus, darbības jomu un vēlākās izmaiņas saistītas ar darbu, kam tās nepieciešamas.

Šeit inženieriem ir noderīga disciplīna. Lieciet ierakstam nest sevī nenoteiktību. Ja avota statuss nav atrisināts, ierakstiet, ka tas nav atrisināts. Ja licence aptver iegūšanu, bet ne apmācību, nesauciet to plaši licencētu. Ja politika mainījās pēc vākšanas, neuzdodiet, ka agrākais lēmums tika pieņemts saskaņā ar vēlāko politiku. Datu kopa nekļūst ticamāka tāpēc, ka tās iezīmes ir pārliecinātas. Tā kļūst labāk pārvaldāma, ja tās iezīmes saglabā to, kas ir zināms, nezināms un nosacīts.

Sarunā jāiekļauj arī izeja

Liela daļa diskusiju pievēršas uzņemšanai: vai pārlūkprogramma šodien drīkst vākt šo avotu? Taču attiecībām ar apmācības datiem ir vajadzīga arī izeja. Licence sasniedz tās beigu datumu. Tiesību īpašnieks maina atrunu. Izdevējs atrod kļūdu atribūcijas ierakstā. Pakalpojumu sniedzējs maina sava modeļa izstrādes plānu. Avots kļūst nepiemērots, jo tā izcelsmi nevar rekonstruēt. Saņemta sūdzība. Tie ir parasti dzīves cikla notikumi, nevis pierādījums, ka kāds ir rīkojies slikti.

Svarīgais jautājums ir, vai sistēma zina, ko darīt, kad tie notiek. Avota politikai būtu jānosaka, kurš saņem paziņojumu, kurš lemj par tā darbības jomu, kurš var apturēt jaunu izmantošanu, kurš var izsekot skartās datu kopas, kurš sazinās ar tiesību īpašnieku un ko organizācija reāli var un ko nevar mainīt jau izlaistā artefaktā. Darba plūsmai jābūt pietiekami konkrētai, lai to varētu pārbaudīt. Teikt, ka bažas tiek uztvertas nopietni, ir pieklājīgs teikums. Tas nav process.

Eiropas diskusijās par modeļu pārvaldību ir paradums atgriezties pie dokumentācijas, jo dokumentācija ir vieta, kur sistēma pasludina savu pašas atmiņu. Tas pats attiecas arī uz šo jomu. Sūdzību izskatīšanas ceļam ir nepieciešams identifikators. Pārskatīšanai ir nepieciešams reģistrēts lēmums. Noņemšanas pieprasījumam ir nepieciešama skaidra robeža. Avota izmaiņām ir nepieciešama versiju vēsture. Modeļa izlaidei jābūt saistītai ar attiecīgajiem apmācības un kūrēšanas ierakstiem. Bez tā pat sirsnīga organizācija galu galā paļaujas uz atmiņu, un atmiņa īpaši labi neiztur pēc dažām personāla maiņām un trim krātuves migrācijām.

Ir arī komerciāls iemesls izeju izveidot agri. Pakalpojumu sniedzējs, kas var izolēt avotu saimi un saprast tās turpmāko izmantošanu, ir ar vairāk iespējām, kad mainās noteikumi. Tas var pārtraukt turpmāko vākšanu, izņemt saimi no plānotās datu kopas, aizstāt ar licencētu materiālu, apturēt izlaidi vai paskaidrot, kāpēc pieprasītais labojums sasniedz vienu artefaktu, bet ne citu. Pakalpojumu sniedzējs bez šādām saitēm ir spiests izteikties plaši, jo tas nevar sniegt precīzu atbildi. Plaši apgalvojumi reti apmierina kādu no pusēm.

Arī tiesību īpašniekiem šajās attiecībās ir pienākumi, lai gan tie nav identiski pakalpojumu sniedzēju pienākumiem. Atruna, kas publicēta skaidrā, stabilā, mašīnlasāmā formā, ir vieglāk ievērojama nekā robeža, kas iestrādāta rindkopā, kuru neviena vākšanas sistēma nevar identificēt. Licencēšanas kontaktpersona, kas var izskaidrot pieejamo ceļu, samazina likumīgas vienošanās izmaksas. Izmaiņu paziņojums, kas saglabā vēsturi, neļauj saprātīgu vācēju vērtēt pēc informācijas, kas tajā laikā nebija pieejama. Slogs nav pilnībā jāpārliek uz izdevējiem, īpaši mazākiem. Vienkārši ir taisnība, ka saskarne darbojas labāk, ja abi gali var runāt.

Komisijas pašreizējais darbs pie atrunu protokoliem tāpēc nav neskaidrs standartu strīds. Tas attiecas uz to, vai tīmeklis var izteikt izvēli tādā mērogā, kādā modeļi vāc materiālu. Labi standarti nenokārtos vērtēšanu, nepadarīs katru juridisko jautājumu vieglu un neizskaudīs ļaunu ticību. Tie var samazināt vienu novēršamas neskaidrības klasi. Tā ir pieticīga ambīcija, un tīmeklī pieticīgas ambīcijas parasti ir tās, kas izdzīvo saskarē ar realitāti.

Ko atbildīgam vācējam būtu jāspēj pateikt

A responsible collector does not need to make legal advice appear in every log line. It does need to make several practical statements true. It should be able to say what purpose a collection served. It should be able to identify the source and version it considered. It should be able to show the access route and the signals it observed. It should be able to identify the rule, licence or unresolved question that governed the decision. It should be able to connect an admitted source to the dataset or system that received it. It should be able to explain how it handles changes and concerns.

Those statements suggest a design rather than a checklist. Begin with a source contract. The contract names the source family, the intended use, accepted access methods, technical limits, known rights signals, required evidence and an owner. It is not a contract in the legal sense unless the parties have made it one. It is an operational contract inside the organisation: a record of what the collection system may do and why.

Then keep decision points close to the actions that matter. Do not fetch first and look for restrictions after the material has been copied through several queues. Do not permit a dataset to move from research evaluation to commercial model training merely because the bytes fit both tasks. Do not transform a licence field into a generic true value because the actual scope is inconvenient. And do not treat a request to stop as a support ticket with no relationship to the source record.

Where a collector cannot understand a signal, it should say so. Where a human review is needed, the system should make the pause visible. Where a licence is available, the commercial and technical routes should meet: the agreement must be represented in a form that changes what the collector can do. A PDF in a legal folder does not prevent a pipeline from doing the wrong thing at three in the morning. The pipeline needs an enforceable state.

At Dweve, the small part of this argument that is directly ours is visible in the training-content material in our Trust Centre. The published record says that its source-family catalogue is not a blanket permission claim and that item, version, licence, attribution, purpose, withdrawal and evidence decisions are controlled in a Spindle ledger. That is a description of our stated process, not a claim that a record resolves copyright questions by itself. Winnow's role in that picture is source intelligence and repeatable, inspectable source processes. The broader point is not about our products. It is that collection becomes more responsible when the source route is treated as evidence rather than as background noise.

This is the kind of product work that rarely appears in a model demo. Nobody applauds a field called observed-terms-version. But those fields decide whether a team can answer a reasonable question without conducting an internal excavation. The future of model governance will include plenty of impressive mathematics. It will also include rather more careful records than the industry once thought fashionable.

The work is governance, not a crawler setting

It is tempting to make this a technical dispute. Put the right instruction in a file, update the user agent, turn a parser on, turn a parser off. Those things matter, but they are only the visible edge of a governance problem. A crawler does what the organisation has decided it may do. If the organisation has no clear purpose boundary, no licence register, no source owner, no review route and no dataset lineage, technical obedience will be inconsistent even when every individual engineer is trying to be careful.

Sāciet ar mērķi, jo mērķis maina jautājumu. Modeli piedāvājošs uzņēmums, kas vāc materiālu šauram, izpaustam pētniecības eksperimentam, ne vienmēr atrodas tādā pašā situācijā kā uzņēmums, kas vāc materiālu komerciālam vispārēja mērķa modelim. Pakalpojumu sniedzējam, kas lūdz licenci izguvei, var būt nepieciešama cita vienošanās nekā tam, kurš vēlas trenēt modeli. Sabiedrības interešu arhīvam var būt citāds pilnvarojums un juridiskais ceļš nekā produktu komandai, kas veido atbilžu dzinēju. Šīs atšķirības nedrīkst izmantot, lai aizēnotu pienākumus. Tās ir iemesls, kāpēc viens universāls statusa lauks nevar atspoguļot patiesību.

Pēc tam padariet īpašumtiesības redzamas. Kādam ir jāpieder avota politikai. Kādam ir jāpieder tehniskajam vācējam. Kādam ir jāpieder licences noteikumiem un lēmumam par avotu saimes pieņemšanu. Kādam ir jāpieder atbildei uz tiesību īpašnieka jautājumu. Vienai personai nav jāuzņemas visas lomas, un maza komanda tās var apvienot. Svarīgi ir tas, ka atbildība ir atklājama, pirms jautājums pārvēršas strīdā. Vispārīga e-pasta kaste nav gluži īpašnieks. Tā ir vieta, kur īpašumtiesības var atrast vai neatrast.

Datu pārvaldībai ir arī jāsatiekas ar modeļa pārvaldību. Lēmums par avotu, kas paliek tīmekļa vākšanas rīkā, kamēr trenēšana notiek citur, pēc būtības ir vājš. Trenēšanas komandai ir jāzina, kuras avotu grupas ir iekļautas tvērumā. Datu kopas veidotājam ir jāsaglabā izņēmumi un nosacījumi. Izlaišanas procesam ir jāzina, vai būtiskas izmaiņas avotu politikā prasa pārskatīšanu. Publiskā kopsavilkuma procesam ir nepieciešams aizstāvams apraksts par trenēšanas satura kategorijām. Nevienam nav jānes visa vēsture savā galvā. Sistēmai ir nepieciešams ceļš, pa kuru vēsturi var izgūt.

Šeit palīdz rūpīga atšķirība starp politiku un izpildi. Autortiesību politika nosaka, ko organizācija plāno darīt un kādiem standartiem tā sekos. Izpilde ir tehnisko un procedurālo kontroļu kopums, kas padara pretēju rīcību grūtāku: piekļuves vārti, avotu ieraksti, pārlūkošanas rīka konfigurācija, piekļuves kontroles, līgumu pārbaudes, datu kopu manifesti, versiju kontrole, pārskatīšana un eskalācija. Politika bez izpildes ir tikai vēlme. Izpilde bez politikas var kļūt par efektīvu veidu, kā īstenot noteikumus, kurus neviens nav pārdomājis. AI akta uzsvars uz autortiesību politiku ir vērtīgs tieši tāpēc, ka tas prasa pakalpojumu sniedzējam savienot juridisko nodomu ar tā darbības praksi.

Nav iemesla, lai tas kļūtu par slepenu iekšēju ceremoniju. Publiska politika var noteikt tvērumu, pieeju tiesību atrunām, izmantoto vākšanas ceļu veidus, kontaktceļu un jebkura publiska apraksta robežas. Tai nevajadzētu solīt noteiktību tur, kur likums vai pierādījumi joprojām ir neskaidri. Tai vajadzētu pateikt, ko tā dara, ja signāls ir neviennozīmīgs, ko tā dara, ja tiesību īpašnieks ar to sazinās, un ko tā var pārskatīt pēc tam, kad avots jau ir izgājis cauri sistēmai. Atturīga politika ir ticamāka nekā vērienīga. Lielākā daļa darba notiek pēc politikas publicēšanas, neizskatīgajā jautājumā par to, vai ieraksts un cauruļvads sakrīt.

Iepirkums ir pelnījis tādu pašu uzmanību. Kad organizācija pērk modeli, datu pakalpojumu vai izguves produktu, tai vajadzētu jautāt vairāk nekā to, vai piegādātājam tīmekļa vietnē ir autortiesību paziņojums. Kādas avotu kategorijas tika izmantotas? Kā tiek apstrādātas tiesību atrunas? Kāda dokumentācija ir pieejama lejupstraumes pakalpojumu sniedzējiem? Vai piegādātājs var izskaidrot savu publisko trenēšanas satura kopsavilkumu? Kas notiek, ja satura avots tiek apstrīdēts, labots vai atsaukts? Kurai pusei ir paredzēts veikt izmeklēšanu? Tie nav jautājumi, ko klients uzdod, lai kļūtu par autortiesību juristu. Tie ir parasti jautājumi par atkarību un pierādījumiem.

The answers may be incomplete, especially in a fast-moving field. Incomplete answers should be labelled as such, with an owner and a route to improve them. The dangerous answer is often the smooth one: all content is public, all training is fair, all records are confidential, all concerns are handled. Each phrase hides the very distinctions that a responsible negotiation needs. A more useful answer names the scope, the rule, the evidence, the limitation and the next review point.

Public authorities have a particular reason to ask these questions. They can influence the market through procurement long before a court or regulator resolves every difficult issue. A tender can require a provider to describe its approach to training data, reservations, licences, provenance and complaints. It can require a record of material changes. It can distinguish a statement of policy from evidence of implementation. It can set proportionate conditions without insisting that the buyer possesses a complete theory of every copyright question in Europe. Procurement is not a substitute for law. It is one of the places where law becomes an operational expectation.

That is the wider lesson. The negotiation is not merely between a publisher and a crawler. It also involves the people who configure collection, create standards, procure models, manage licences, build datasets, publish model documentation, operate complaint routes and decide whether a model can be placed on the market. A training-data relationship becomes less extractive when those people can see the same evidence and work with the same boundaries. The alternative is not simplicity. It is a chain of separate assumptions that eventually meets a rightsholder at the worst possible moment.

Public does not mean ownerless

The public web should remain public. That means it should remain possible to read, link, quote within the law, research, index, criticise, preserve and build new services. A web made of permission pop-ups for every ordinary act would not be open in any useful sense. But openness is not the same thing as ownerlessness, and visibility is not an all-purpose transfer of rights.

The negotiations now forming around training data are an opportunity to make that distinction practical. A rightsholder should be able to reserve a use in a form a responsible collector can receive. A model provider should be able to obtain high-quality material through terms that can be implemented and audited. Small organisations should not be excluded because only the largest parties can afford bespoke agreements. Public authorities should be able to see enough about model-training content and copyright policy to ask intelligent questions. And when a dispute or change occurs, both sides should have a route that begins with evidence rather than theatre.

There will still be hard cases. The law will still need interpretation. Some rights conflicts will not be fixed by a standard, and some commercial negotiations will remain unequal. A provenance ledger will not make a poor bargain fair. A technical marker will not make a right meaningful if a collector deliberately ignores it. The aim is not a frictionless web. Frictionless often means that somebody else's cost has been hidden.

The better aim is an inspectable web. One in which an automated request can carry a declared purpose, a published boundary can survive contact with a collection system, a licence can become an operational rule, and a model provider can account for the route that brought material into its work. That is a less romantic vision than a machine reading everything ever written. It is also more likely to leave publishers, creators, researchers and model builders with a web worth negotiating over.

Sources