Efficient and accurate systems for querying unstructured data
Abstract: (shortened) In recent years, automatic analysis over this unstructured data has become possible via machine learning (ML). Analysts can use ML to extract structured information from these unstructured sources, such as object types and location from a video. The structured information can subsequently be used in downstream analysis, e.g., the urban planner can count the number of cars that passed by an intersection. Unfortunately, using ML for these analyses is challenging. Deploying ML is prohibitively expensive for many organizations: naively analyzing a year of video from a small town can cost millions in cloud compute credits. ML methods are also unreliable, returning incorrect results, which can lead to downstream errors. Finally, deploying ML for analytics requires knowledge of deep learning, data systems, programming, and other technical skills. In light of these challenges, we make two observations: many applications can tolerate approximations, if there are guarantees on accuracy, and methods for answering unstructured data queries range by up to 10 orders of magnitude in cost. In this dissertation, we develop systems and algorithms for efficient and reliable unstructured data analytics, leveraging the two observations. Instead of returning exact answers, we return approximate answers generated by cheap approximations to expensive ML methods. Our systems can return statistically valid answers on a wide range of query types, including selection, aggregation, and limit queries. Furthermore, our systems can be up to orders of magnitude cheaper than standard methods of answering queries Comments
Source attribution: Lobsters AI. Reader content is derived from the canonical public URL when extraction is available.
Reader mode
Status: not requested
Reader extraction has not been requested for this article.
AI reading tools
Usable reader text is required before AI tools can run for Efficient and accurate systems for querying unstructured data.
Related articles
5 recommendations
- Hacker News AI
Beyond Context Windows: Evaluating Long-Term Memory for AI Agents
AI agentsContext engineeringModel evaluationRead related article: Beyond Context Windows: Evaluating Long-Term Memory for AI Agents - Lobsters AI
LLMs Are Too Big. My Log Router Doesn't Need to Sing
Model evaluationRead related article: LLMs Are Too Big. My Log Router Doesn't Need to Sing - OpenAI Blog
Building standards for the next phase of AI
OpenAIModel evaluationRead related article: Building standards for the next phase of AI - Google DeepMind
Piloting the world's first double-blind AI evaluations
Google DeepMindModel evaluationRead related article: Piloting the world's first double-blind AI evaluations - Hacker News AI
Show HN: An AI office-work benchmark, and 4 bugs we found in our own judge
AI agentsModel evaluationRead related article: Show HN: An AI office-work benchmark, and 4 bugs we found in our own judge