Unlocking Unstructured Data in the Public Sector
State, local, and education (SLED) organizations have been advancing for decades: Most public sector data strategies are built around tables and scheduled reports. That made sense for a long time. It’s what the job required. But agencies also hold an enormous amount of data that…
Share this post:
State, local, and education (SLED) organizations have been advancing for decades: Most public sector data strategies are built around tables and scheduled reports. That made sense for a long time. It’s what the job required.
But agencies also hold an enormous amount of data that doesn’t fit neatly into rows and columns: PDFs and scanned forms, audio and call transcripts, images, video, logs, and sensor feeds.
Analysts estimate that 80–90% of newly generated enterprise data is unstructured, and public sector organizations are no exception. The agencies that figure out how to use it won’t just have more data. They’ll have better context for the decisions they’re already making.
Why unstructured data matters
Unstructured data not only captures nuance but also contains information that simply can’t be expressed in tables – conversations, images, environmental signals, video, and free-form text. And it already connects to core SLED missions.
Here’s just a few examples of how some public sector organizations are putting unstructured data to work today:
- Public safety: body-worn camera video, license plate images, and 911 call audio to accelerate investigations and improve transparency.
- Health & human services: caseworker notes, scanned applications, and provider documentation to spot eligibility issues or speed benefits decisions.
- Education: student essays, lecture recordings, and transcripts to support tutoring, accessibility, and early-warning indicators.
- Transportation & infrastructure: drone imagery, right-of-way photos, and sensor streams for maintenance and safety.
- Courts & legal: scanned filings and hearing audio to improve search, disclosure, and public access.
- Environmental & emergency response: satellite imagery, weather models, and incident reports to inform planning and response.
Why it’s been hard (and why that’s changing)
Government data environments were designed around structured data. The result is a patchwork of tools and storage patterns that makes sharing new data types difficult. Add legitimate compliance responsibilities, like HIPAA, CJIS, FERPA, and it’s easy to see why teams have been cautious when governance models aren’t clear.
But that tradeoff has shifted. Modern approaches let agencies manage unstructured and structured data together with policy-driven controls, lineage, auditability, and role-appropriate access. The choice is no longer between using the data and protecting it. And because modern approaches maintain lineage across both structured and unstructured sources, agencies can answer the harder question too. Not just what the data showed, but where it came from and how it informed a decision.
What “starting” can look like
Every organization’s path is different, but early wins often come from:
- Document understanding: extracting key fields from PDFs and scanned forms (claims, applications, permits).
- Audio/text analysis: summarizing call transcripts and case notes to surface themes, risks, or follow-ups.
- Image/video workflows: classifying, tagging, or redacting sensitive content to speed reviews and improve privacy.
- Sensor and log enrichment: correlating events across systems to reduce outages and improve response times.
None of this requires an overnight rebuild. Tacking simple, well-governed, and high-value use cases help agencies learn quickly and prove impact.
As technology evolves, so does the opportunity to use data more fully. Bringing unstructured information into your strategy is one way to keep modernization moving at a pace that works for your team and the communities you serve.
To see where unstructured data is getting stuck in your current environment, use The Data Unification Diagnostic.
Last updated: December 9, 2025
The first unified platform to bring the power of AI to your data and people, so you can deliver AI’s potential to every constituent.
Databricks is a leading data and artificial intelligence (AI) company, founded by the original creators of Apache Spark™, Delta Lake, and MLflow. Their mission is to simplify and democratize data and AI so that every organization can harness its full potential.