Research Workspace

Our research focuses on how models read documents; retrieval and embedding strategies, vision-language models, and Document VQA. It starts with problems we hit building real pipelines, and we publish our findings in full, including negative results.

Vision-Language Models

Our work on vision-language models examines how they perform on document inputs rather than natural images, with a focus on visual grounding and sensitivity to page layout.

Document VQA

Document VQA covers question answering over forms, invoices, tables, and scanned material, and the failure modes that appear as document quality and structure vary.

Retrieval and RAG

Our retrieval work evaluates chunking, reranking, and multimodal retrieval on real document collections to establish where retrieval improves accuracy and where it does not.

Meet our team

Muhammad Talal Majeed

Researcher

Muhammad Talal Majeed

Hi! I'm Talal, a tech student in Islamabad interested in closing the gap between the software industry and academia. I started with game development in 2018, moved through web development, and now work part-time in MLOps and DevOps. In 2026 I founded ecello.net, a place for collaborative research. My research sits in deep learning for document intelligence.

Momena Akhtar

Researcher

Momena Akhtar

Hi! I'm Momena, a final-year Computer Science student at NUST. I'm currently researching Document VQA at DFKI, and on the product side I build and ship web and cloud applications. I like working across both: research questions that come out of real systems, and systems that get better because of what that research turns up. That mix is what keeps me interested.

Our blog

Coming soon

We are still writing up our first results. Papers and notes will be listed here as they go out.

Get in touch

If you are working on document understanding, retrieval, or vision-language models and want to compare notes, collaborate, or reproduce something we published, write to us.