pikuri-pdf
pikuri-pdf plugs PDF → text extraction into pikuri-core's +Pikuri::Extractor+ registry. The bundled +Pikuri::Extractors::PDF+ extractor wraps the pure-Ruby pdf-reader gem and extracts lazily: paged reads (the +read+ tool's windows) parse only the pages the window needs, so the first page of a 500-page PDF never pays for the other 499. Shipped separately from pikuri-core so the core's dependency tree stays minimal and auditable: pdf-reader and its transitive deps (Ascii85, afm, hashery, ruby-rc4, ttfunk) ride along only for hosts that opt into PDF support. Registration is explicit — +Pikuri::Extractors::PDF.register+ — so requiring the gem changes nothing by itself; the host script picks which extractors it wires in. One registration extends the +read+ tool, +web_scrape+, and the pikuri-vectordb indexer simultaneously.
Activity
- Latest release
- 2w ago
- Total releases
- 3
- Cadence
- ~43 days
- Last 12 months
- 3
Reach
- Downloads
- 603
Details
- License
- MIT
- First release
- Jun 04, 2026
| Version | Released | |
|---|---|---|
0.1.0
minor
|
0.1.0
minor
Dependencies (2)
|
|
0.0.7
patch
|
0.0.7
patch
Dependencies (2)
|
|
0.0.6
initial
|
0.0.6
initial
Dependencies (2)
|