Kymata Labs/The Living IndexesBuilt by tekvisions ↗
The Document Index / PDF Extraction / #121
CatchTheTornado

CatchTheTornado/text-extract-api

by CatchTheTornado · PDF Extraction · updated 9mo ago

Document (PDF, Word, PPTX ...) extraction and parse API using state of the art modern OCRs + Ollama supported models. Anonymize documents. Remove PII. Convert any document or picture to structured JSON or Markdown

39
momentum
3,176
stars
279
forks
#121
rank
anonymizationapiextractjsonllmocrocr-pythonpdfpii
View on GitHub →