Fatma

Fatma Özcan is a Principal Software Engineer at Google. Prior to that, she was a Distinguished Research Staff Member and a senior manager at IBM Almaden Research Center. Her current research focuses on platforms and infra-structure for large-scale data analysis, knowledge graphs, democratizing analytics via NLQ and conversational interfaces to data, and query processing and optimization of semi-structured data. Dr Özcan got her PhD degree in computer science from University of Maryland, College Park, and her BSc degree in computer engineering from METU, Ankara. She has over 20 years of experience in industrial research, and has delivered core technologies into many IBM products. She is the co-author of the book "Heterogeneous Agent Systems", and co-author of several conference papers and patents. She is an elected member of the SIGMOD Executive Committee, and is on the board of trustees for the VLDB Endowment and Computing Research Association. She is an ACM Distinguished Member.
Authored Publications
Sort By
  • Title
  • Title, descending
  • Year
  • Year, descending
Fine-Grained Table Retrieval for Open-Domain Tabular Question Answering
Xingyu Ji
Wojciech Kosiuk
Madelon Hulsebos
Proceedings of the 11th Workshop on Automated Knowledge Base Construction (AKBC 2026), Association for Computational Linguistics
Preview abstract This work introduces a fine-grained table retrieval framework for grounding large language models in heterogeneous, open-domain relational data. Instead of encoding a query as a single vector, the approach decomposes natural language queries into semantic components and embeds each independently, enabling more precise matching of compositional query intent. These representations are used in a staged retrieval pipeline with component-level search, connectivity-aware grouping, and reranking. Experiments on three TARGET benchmark corpora show consistent improvements in capped recall@k and stronger alignment between query intent and tabular structure over dense retrieval baselines, particularly for longer and more complex queries and when using lightweight embedding models. View details
SemBench: A Benchmark for Semantic Query Processing Engines
Jiale Lao
Gerardo Vitagliano
Immanuel Trummer
H. V. Jagadish
Sebastian Schelter
Andreas Kipf
Matthew Russo
Kris Kissel
Michael Cochez
Andreas Zimmerer
Olga Ovcharenko
Thibaud Hottelier
Gautam Gupta
Tianji Cong
2026
Preview abstract We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on zero-shot abilities of state-of-the-art large language models (LLMs). They extend SQL with semantic operators, configured by natural language instructions, that are evaluated via LLMs and enable users to perform various operations on multimodal data. Our benchmark provides variety along three axis: scenarios, modalities, and operators. Included are scenarios ranging from movie review analysis to medical question-answering. Within these scenarios, we cover different data modalities, including images, audio, and text. Finally, the queries involve a diverse set of operators, including semantic filters, joins, mappings, ranking, and classification operators. We evaluate systems according to processing overheads and result quality. We present experimental results for an industrial semantic query processing engine (BigQuery), as well as academic systems (LOTUS, Palimpzest, and ThalamusDB). Our results shed light on the relative strengths and weaknesses of the evaluated systems, and hint at promising avenues for future research. View details
Preview abstract The integration of vector search into databases, driven by advancements in embedding models, semantic search, and Retrieval-Augmented Generation (RAG), enables powerful combined querying of structured and unstructured data. This paper focuses on filtered vector search (FVS), a core operation where relational predicates restrict the dataset before or during the vector similarity search (top-k). While approximate near neighbor (ANN) indices are commonly used to accelerate vector search by trading latency for recall, the addition of filters complicates performance optimization and makes achieving stable, declarative recall guarantees challenging. Filters alter the effective dataset size and distribution, impacting the search effort required. We discuss the primary FVS execution strategies – pre-filtering, post-filtering, and inline-filtering – whose efficiencies depend on factors like filter selectivity, cardinality, and data correlation. We review existing approaches that modify index structures and search algorithms (e.g., iterative post-filtering, filter-aware index traversal) to enhance FVS performance. This tutorial provides a comprehensive overview of filtered vector search, discussing its use cases, classifying current solutions and their trade-offs, and highlighting crucial research challenges and future directions for developing efficient and accurate FVS systems.   View details
Preview abstract Large Language Models (LLMs) have demonstrated impressive capabilities across a range of natural language processing tasks. In particular, improvements in reasoning abilities and the expansion of context windows have opened new avenues for leveraging these powerful models. NL2SQL is challenging in that the natural language question is inherently ambiguous, while the SQL generation requires a precise understanding of complex data schema and semantics. One approach to this semantic ambiguous problem is to provide more and sufficient contextual information. In this work, we explore the performance and the latency trade-offs of the extended context window (a.k.a., long context) offered by Google's state-of-the-art LLM (\textit{gemini-1.5-pro}). We study the impact of various contextual information, including column example values, question and SQL query pairs, user-provided hints, SQL documentation, and schema. To the best of our knowledge, this is the first work to study how the extended context window and extra contextual information can help NL2SQL generation with respect to both accuracy and latency cost. We show that long context LLMs are robust and do not get lost in the extended contextual information. Additionally, our long-context NL2SQL pipeline based on Google's \textit{gemini-pro-1.5} achieve a strong performance with 67.41\% on BIRD benchmark (dev) without finetuning and expensive self-consistency based techniques. View details
×