Modern organisations increasingly store valuable information as unstructured text: reports, support tickets, policies, scientific notes, web content, and AI-generated reasoning traces. While databases provide declarative operations for structured data—such as INSERT, UPDATE, DELETE, and SELECT—manipulating unstructured data still usually requires manually writing document-by-document Large Language Model(LLM) workflows.
This project explores DocDML, a system for declarative manipulation of unstructured data with LLMs. Instead of writing a separate procedure for each document, users specify what content should be selected, how it should be changed, and what conditions must remain true after the change.
For example, a user may ask the system to find explanations that are unnecessarily confusing, replace only the unclear part with a simpler explanation, and keep the final conclusion unchanged. The system identifies the relevant text, makes a small targeted update, checks that the result is still correct, and applies the change efficiently across a large collection.
The student will contribute to the DocDML prototype by designing and evaluating declarative operators for selecting, updating, deleting, and querying textual data. The work will investigate how LLMs can act as semantic predicates, patch generators, and validators, while database techniques provide batching, versioning, provenance, and efficient execution.
Computer Science and Engineering
Database systems | Data management | Large language models | Artificial intelligence | Natural language processing
No
- Research environment
- Expected outcomes
- Supervisory team
- Reference material/links
The student will work within the DKR group at UNSW Computer Science and Engineering, supported through regular meetings with the supervisory team. The project provides hands-on experience in database systems, language-model applications, experimental evaluation, and research software development.
The student will have access to relevant computing infrastructure, language models, research datasets, and existing open-source declarative AI systems. Experience with Python is desirable; prior research experience is not required.
- New declarative operators for manipulating unstructured textual chunks.
- A prototype for semantic selection and local text updates using LLMs.
- Support for constraints such as preserving specified content, maintaining verified answers, and limiting the size of an update.
- A batching and caching strategy for applying one high-level observation across many textual records efficiently.
- An evaluation comparing local semantic updates with full-document LLM rewriting, including correctness, update locality, model cost, and throughput.
- Reproducible implementation, technical report, and poster/demo.