Tensorlake

@tensorlake.ai

The tensorlake playground was, unlike AWS textract and every other tool I have tried, able to parse my angled, low-quality scan of Norwegian pay statistics from 1926. Not that 1926 Norwegian statistical tables is a generally useful benchmark…

Tensorlake@tensorlake.ai · 9mo ago

Document parsing benchmarks have been measuring the wrong thing. We tested every major parser on real enterprise documents. The results will change how you think about OCR accuracy 🧵

Two dense document pages flank a skeptical person’s sticker-style portrait against a green gradient, link text centered below.

Document parsing benchmarks have been measuring the wrong thing. We tested every major parser on real enterprise documents. The results will change how you think about OCR accuracy 🧵

Two dense document pages flank a skeptical person’s sticker-style portrait against a green gradient, link text centered below.

New: Vision Language Models now power key document processing features We're using VLMs for: - Page classification in large documents - Table/figure summarization - Fast structured extraction (skip_ocr mode) Here's what this means for document processing 🧵

Most parsers strip all tracked changes when you extract the text. That means: ❌ Lost audit trails ❌ Manual review of revision history ❌ No programmatic access to reviewer comments ❌ Workflows that can't route based on specific edits

Tensorlake interface showing parsed Word document with tracked changes preserved as HTML tags, displaying an insurance claim report

OCR engines constantly mess up document hierarchy. Section 2.2 becomes a top-level header (##) instead of nested (###). We just shipped automatic header correction. 🧵 How it works:

Comparison of document header detection. Left side "Just OCR" shows incorrect hierarchy with section 2.2 at wrong indent level. Right side "Header Correction" shows proper nesting where 2.2 is correctly indented under section 2. Bottom shows Python code: doc_ai.parse_and_wait() with cross_page_header_detection=True parameter. Green gradient background with Tensorlake logo.

Citations. When users ask "where did this come from?" your system should point to the exact page fragment...not just "file_name.pdf". Built citation-aware RAG with spatial metadata has: → Parse docs with bounding boxes → Embed citation anchors in chunks → Return page numbers + coordinates A 🧵

In finance, clinical trials, or performance benchmarks, dense tables contain mission-critical data. But flatten that data like most parsers do and trust is lost. Tensorlake restores trust by preserving structure, generating summaries for effective embeddings, and attaching evidence via b-boxes.

Side-by-side comparison of a dense healthcare data table in PDF format and its structured DataFrame output. A green background with the Tensorlake logo shows an arrow pointing from the PDF to the DataFrame. The caption reads “Parse Dense Tables Reliably” with the link “tlake.link/blog/dense-tables” at the bottom.

You can now login into Tensorlake using Microsoft and Azure SSO credentials. This is the beginning of better integration with Microsoft Azure and Tensorlake. If you are using Azure, and need better Document Ingestion and ETL for unstructured data reach out to us!

Bild

“RAG is dead” is lazy. What’s dead is cosine‑N without a retrieval plan. We ship advanced RAG...out of the box: • Classify pages → target sections • Extract structured fields → filter by form_type, fiscal_period • Verify data; cite page/bbox Want to know how? 🧵👇

Screenshot of a Tensorlake blog post titled Accelerate Advanced RAG with Tensorlake by Dr. Sarah Guthals, dated August 19, 2025. The header image reads ‘RAG isn’t dead, undisciplined retrieval is’ with a green wave design and Tensorlake logo. The page shows the TL;DR summary emphasizing the need for a retrieval plan over naive Top-N cosine RAG, and a table of contents with sections on the Freshness Principle, Accelerate Advanced RAG, a real-world application fact-checking Tesla SEC filings, and treating context as a hard requirement.

Build a smart real estate agent (no license required). 🧠 LangGraph (by @langchain.bsky.social) + 📝 Tensorlake Contextual Signature Detection = ✅ Knows who signed ✅ When they signed ✅ If it’s ready to close Full tutorial + code linked below 👇

Most "unstructured" parses fail on when layout gets tricky: multiple columns, fragmented text blocks, mixed reading order Tensorlake doesn't. ✅ Authors parsed as one clean chunk ✅ Abstract follows, exactly as it should Unstructured ≠ unordered Preserve reading order. Parse with Tensorlake.

BildBild

Over the last few weeks we have been working a ton on some huge improvements to the @tensorlake.ai API and SDK. They are finally live 🥳 More announcements around this is coming soon, but if you didn't see the announcement in our Slack, make sure you use v2 API and SDK 0.2.20 🙌

Want to see the tool in action? Check out this quick demo or try it out in the Colab Notebook (linked in the comments)

LangChain + Tensorlake: Unlocking Document Understanding for Agents

YouTube video by Tensorlake

youtube.com

Tensorlake@tensorlake.ai · last yr.

Just published 🐦‍⬛ langchain-tensorlake 💚 A new @langchain.bsky.social tool to parse real-world documents (PDFs, scans, forms) with Tensorlake & feed structured data right into your agents. Built for devs wrangling docs in legal, finance, healthcare & more. Learn more: tlake.link/langchain-tool

Build a smart real estate agent (no license required). 🧠 LangGraph (by @langchain.bsky.social) + 📝 Tensorlake Contextual Signature Detection = ✅ Knows who signed ✅ When they signed ✅ If it’s ready to close Full tutorial + code linked below 👇

21 hours later and we’re in the top 5 on Product Hunt! 🚀 Huge thanks to everyone who supported, upvoted, and shared 💚 Tensorlake is just getting started. Stay tuned - there’s so much more to come. P.S. There's still time to upvote our launch and let us know your thoughts 👇 #AI #RAG #LLM #devtools

Tensorlake - Parse documents like a human & build Python-based workflows | Product Hunt

Tensorlake Cloud is a platform for document ingestion and data orchestration. Parse real-world documents with human-like layout understanding and build Python-based workflows at scale and ready for pr...

producthunt.com

🚀 We’re in the top 3 on Product Hunt today just 6.5 hours after launch! Huge thanks to everyone supporting Tensorlake 🎉 From devs wrangling PDFs to teams automating high-stakes workflows. If you haven’t yet, check us out 👇

Bright green graphic with the Product Hunt logo featuring a cat wearing Google Glass at the top. Below it, large bold black text reads: “We Are Live on Product Hunt.” At the bottom of the image is the Tensorlake logo and name in white.