The tensorlake playground was, unlike AWS textract and every other tool I have tried, able to parse my angled, low-quality scan of Norwegian pay statistics from 1926. Not that 1926 Norwegian statistical tables is a generally useful benchmark…
Document parsing benchmarks have been measuring the wrong thing. We tested every major parser on real enterprise documents. The results will change how you think about OCR accuracy 🧵