Saišu grāmata un dokumenti bez līmēšanas koda

Dokumenti izskatās kā faili, līdz īstai cauruļvadam tie jāatver. Bindery pārvērš biroja formātus vienā Rust dzinējā ar atpazīšanu, normalizētām API, DocQL...

Saišu grāmata un dokumenti bez līmēšanas koda

The file was not the problem

The first mistake is thinking the document is the file. It arrives as a file, yes. It has a name, usually with an extension that looks reassuring. The procurement spreadsheet says .xlsx, the report says .docx, the archive dump says .pdf, and everyone pretends the world has become simple because the last four characters look familiar. Lovely. Then the extension lies, the workbook has hidden sheets, the slide deck contains embedded objects, the PDF is mostly text but not quite, and the old Word file still smells faintly of 2003.

Most teams do not build a document pipeline. They build a small museum of format exceptions. One library for Word, another for Excel, something else for PDFs, a heroic shell script for the archive folder, a Python package that was last updated when everyone still thought QR codes were exciting, and a few regular expressions that should be taken outside and given a quiet retirement. Six months later the glue layer is larger than the product. This is not a rare failure mode. This is the normal shape of document work when every format gets its own kingdom.

Bindery exists because the document layer should not become the main project. At the implementation level it isa Rust crate named dweve-bindery. The public page describes one crate with 17 formats, DocQL, Python bindings, and a common engine. The important part is that high-level APIs like Document, Presentation, and Workbook sit over format-specific modules for OLE2, OOXML, ODF, iWork, RTF, PDF, Markdown, EPUB, LaTeX, images, formulas, and DocQL. That list is not there to impress anyone. It is there because real corpora are rude.

The useful claim is simple: open the document through one engine, normalize what can be normalized, and keep the format-specific pain below a shared surface. That does not make every format identical. It makes the differences explicit enough that a pipeline can survive them.

Bindery starts by distrusting the extension. Detection, parsing, normalization, querying, and writing are separate jobs, which is exactly why the glue does not have to leak everywhere.

Detection is not a decoration

Format detection sounds like a small utility until it fails. Then it becomes the whole incident. Extensions are metadata supplied by the person, tool, mail gateway, export job, migration script, or tired intern that last touched the file. Sometimes they are right. Sometimes they are a polite suggestion. A serious document engine should inspect magic bytes, container structure, package parts, streams, and internal cues before deciding which reader owns the file.

Bindery treats that as the front door. The README and page both describe automatic format detection. The library documentation shows Document::open and Presentation::open as the normal path, not a choose-your-parser ceremony. That matters because users do not want a training course in document archaeology before they can extract a table. They want the engine to pick the path and then give them a stable API.

There is a dry little lesson here. The less glamour a component has, the more damage it can do when people wave it away. Detection is notglamorous. Neither is encoding conversion, ZIP handling, OLE directory walking, relationship resolution, Snappy decompression, or XML namespace handling. Fine. The boring work is exactly where production pipelines either become reliable or begin collecting lucky charms.

One model does not mean one lie

A unified API can become dangerous when it pretends differences have vanished. Bindery should not claim that a PDF, a spreadsheet, an iWork archive, and an OOXML package are the same animal wearing different hats. They are not. The useful architecture is not to flatten truth into mush. It is to expose common operations where they are common and keep capability boundaries visible where they are not.

The source layout shows that split. There is a unified Word document API, a unified presentation API, spreadsheet traits, formula evaluation behind features, DocQL connectors, and lower-level modules for the formats themselves. The public format matrix says OOXML and ODF are first-class read, write, and query surfaces. PDF and RTF are read-heavy and more careful on writing. EPUB, LaTeX, and Markdown are output formats. Legacy Office and iWork have their own internal machinery. That is the right posture. One engine, yes. One fantasy, no.

This distinction matters in audits and data products. If a compliance workflow extracts clauses from contracts, it must know whether a value came from a paragraph, a table cell, a slide note, a formula, or a PDF text run. If an ingestion job feeds retrieval, it must know whether images, comments, metadata, and relationships were preserved, ignored, or marked as unsupported. The answer cannot be buried inside a parser-specific footnote, because that footnote will not show up when someone asks why the result changed.

A shared engine still needs a capability matrix. The honest promise is not that every format behaves the same. It is that each promise is named and testable.

Documents are structured data that forgot to admit it

The worst thing a document pipeline can do is reduce everything to text too early. Text is useful. Text is not the whole document. A spreadsheet has formulas, references, sheets, rows, cells, number formats, comments, and workbook structure. A presentation has slides, shapes, images, notes, ordering, and sometimes a corporate template that has survived three mergers and one rebrand through sheer malice. A Word document has paragraphs, runs, tables, headers, footers, styles, relationships, and embedded objects. A PDF has streams and layout decisions that may or may not correspond to reading order. Turning all of that into one flat string is fast, comforting, and often wrong.

Bindery’s high-level APIs are useful because they keep document shape alive long enough to ask better questions. The document API exposes paragraphs, runs, tables, rows, and cells. The spreadsheet module exposes workbook and worksheet traits. The formula engine covers a large Excel-compatible function surface. DocQL adds a SQL-like query language over the document model, with lexer, parser, validator, planner, executor, connectors, values, and functions in the source tree. That is more than a convenience wrapper. It is a way to stop rewriting the same extraction logic for every format.

Imagine asking one corpus question: which cells reference this assumption, which tables contain a risk category, which slides mention a policy, which documents have a custom property, and which formulas depend on a given input. In the per-format kingdom, that becomes four scripts and a spreadsheet of apologies. In a shared model, it becomes a query surface. Still work, obviously. Software rarely gifts you a holiday. But it is the right work.

DocQL ir atšķirība starp teksta nokasīšanu un dokumentveida jautājumu uzdošanu. Atsauces, formulas, tabulas, formas un metadati paliek daļa no darba.

Kāpēc Rust ir saprātīga vieta šim juceklim

Dokumentu formāti ir brīnišķīga kombinācija no binārām struktūrām, saspiestām pakotnēm, XML, mantotiem kodējumiem, attēlu datiem, datumu sistēmām, formulu semantikas un drošības apsvērumiem. Citiem vārdiem, tāds darbs, kur izplūdusi atmiņas pārvaldība ir dzīvesstila izvēle ar rēķiniem klāt. Rust ir saprātīga bāze, jo Bindery ir jāveic rūpīga parsēšana, jāpārvalda buferi, skaļi jāapstrādā kļūdas un jāpiedāvā API, kas neliek pārējam stekam minēt, kas nogāja greizi.

Crate funkcijas stāsta to pašu stāstu. Noklusējuma funkcijas ietver OLE, OOXML, OOXML šifrēšanu un eval dzinēju. Pilns atbalsts ieslēdz iWork, ODF, RTF, formulas, attēlu konversiju, fontus un vēl vairāk. DocQL ir funkcija ar savu bināro failu. Neobligātās atkarības sakrīt ar formātiem un virsmām, ko tās atbalsta: ZIP apstrāde, ātra XML parsēšana, kodējumu konversija, Snappy, protobuf, attēlu dekodēšana, statistika un kompleksie skaitļi formulu darbam utt. Tas nav viens milzu blobs, kas izliekas, ka katra atkarība pieder visur. Funkciju karodziņi saglabā dokumentu dzinēja formu redzamu.

Tas ir svarīgi iegulšanai. Zināšanu sistēma var vēlēties pilnu biroja un vaicājumu virsmu. Neliels pakalpojums var vēlēties tikai OOXML un teksta izvilkšanu. Python darbplūsma var vēlēties saistījumus virs tā paša dzinēja. Komandrindas pārbaudes ceļš var noderēt vienreizējiem vaicājumiem un testiem. Lapa runā par Rust, PyO3 un CLI virsmu; avots rāda DocQL bināro failu un PyO3 pakotni. Svarīgā dizaina izvēle ir tā, ka šie ieejas punkti balstās uz vienu dzinēju. Citādi katra integrācija kļūst par savu nedaudz atšķirīgu patiesību, un tad kļūdu ziņojumi sāk nēsāt dažādas cepures.

Ieejas punkti var atšķirties bez patiesības sašķelšanas. Rust API, Python saistījumi, DocQL un plašāks Dweve stekam vajadzētu patērēt to pašu dzinēju.

Uzturēšanas stāsts ir produkta stāsts

Bindery ir viegli aprakstīt kā parsētāju, bet uzturēšanas stāsts ir īstais produkta stāsts. Katrai jaunai formātu bibliotēkai, kas pievienota cauruļvadam, ir savs izlaišanas ritms, kļūdu vārdnīca, kļūdu tipi, dīvainības, atkarību risks, testa dati un atteices režīmi. Mazā mērogā tas izskatās pārvaldāmi. Korpusa mērogā tas kļūst par operacionālu dokumentu skapi, kas kož.

Kopīgs dzinējs nenoņem formātu sarežģītību. Tas būtu aizdomīgi. Tas pārvieto sarežģītību vietā, kur testus, iespēju etiķetes, funkcijas un API var pārvaldīt kopā. README ir beigu gala testi dokumentiem, prezentācijām, izklājlapām, iWork un citiem formātiem. Avota koks satur moduļus, kas padara formātu robežas acīmredzamas. Šī struktūra ļauj komandai uzlabot parsētāja slāni bez nepieciešamības katrai produkta komandai no jauna mācīties atšķirību starp attiecību daļu un saliktu failu straumi.

Lūk, kāpēc Bindery sader ar pārējo Dweve kopuma daļu. Reed rūpējas par parsēšanu ar čekiem. BitWeave rūpējas par deterministisku izguvi. Fabric un Spindle rūpējas par pārvaldītām zināšanām un operatīvu lietojumu. Bindery atrodas pirms šiem slāņiem. Tas pārvērš biroja dokumentus strukturētā materiālā, par ko pārējā kopuma daļa var spriest. Ja ievades slānis ir līme un cerība, lejupstraumes sistēma manto līmi un cerību. Ļoti efektīvi, ja jūsu stratēģiskais mērķis ir nākotnes ciešanas.

Kas jājautā pirms tā ieviešanas

Pirmais jautājums nav tas, vai Bindery atbalsta jūsu iecienītāko paplašinājumu. Tas ir kontroljautājums, un kontrolsaraksti ir labi, bet ar to nepietiek. Labāks jautājums ir, kuri dokumentu solījumi jums ir nepieciešami. Vai jums ir nepieciešama tikai lasāma ekstrakcija, rakstīšanas atbalsts, pilna cikla saglabāšana, formulu aprēķināšana, vaicājumi pāri formātiem, metadati, iegulti attēli, šifrēts OOXML, mantots Office, iWork, ODF vai izvade uz EPUB, LaTeX un Markdown? Tie ir dažādi darbi. Izlikšanās, ka tie ir viens darbs, ir tas, kā ceļveži pārvēršas zupā.

Otrais jautājums ir, kā tiek atklātas kļūmes. Ja parsētājs nevar saglabāt struktūru, vai tas to pasaka? Ja rakstītājs ir labāko centienu līmenī, vai tas ir redzams? Ja formulu nevar aprēķināt, vai izsaucējs var izlemt, vai bloķēt, brīdināt vai turpināt? Dokumentu dzinējs nav uzticams tāpēc, ka tas nekad nesaka nē. Tas ir uzticams tāpēc, ka tā nē ir tipizēts, konkrēts un tuvu problēmai.

Trešais jautājums ir, kā jūs testējat savu korpusu. Publiski piemēri ir noderīgi, bet jūsu arhīvs, iespējams, ir dīvaināks par piemēru mapi. Tajā ir bojāti eksporti, sena veidnes, dīvaini skaitļu formāti, paslēptas lapas, kopēti-ielīmēti tabulas, nekorekti PDF un faili ar nosaukumu final_final_really_final. Bindery nodrošina kopīgu dzinēju un testējamas virsmas. Jums joprojām ir nepieciešami korpusa testi. Diemžēl dokumenti paši nekļūs pieauguši.

Mācība

Bindery mācība ir tāda, ka dokumentu apstrāde nav teksta ekstrakcija ar papildu soļiem. Tā ir formāta noteikšana, strukturāla parsēšana, spēju robežas, vaicājumu virsmas, rakstīšanas ceļi un uzturēšanas disciplīna. Lietotājs redz failu. Sistēma redz konteineru, straumes, attiecības, ierakstus, stilus, formulas, metadatus, kodējumus un izvades solījumus. Labs dzinējs notur šo sarežģītību zem produkta virsmas, neizliekoties, ka tās nav.

Bindery uzdevums ir padarīt dokumentu slāni garlaicīgu labā nozīmē. Viens Rust dzinējs. Augsta līmeņa API dokumentiem, prezentācijām un izklājlapām. Formātu moduļi sarežģītajām daļām. DocQL kopīgiem jautājumiem. Python un kopuma integrācija, kur tas ir noderīgi. Spēju etiķetes formātu mitoloģijas vietā.

Fails nekad nebija problēma. Līme bija. Bindery ir mēģinājums pārtraukt maksāt īri par šo līmes slāni.