Vision-Language Models
Our work on vision-language models examines how they perform on document inputs rather than natural images, with a focus on visual grounding and sensitivity to page layout.
Document VQA
Document VQA covers question answering over forms, invoices, tables, and scanned material, and the failure modes that appear as document quality and structure vary.
Retrieval and RAG
Our retrieval work evaluates chunking, reranking, and multimodal retrieval on real document collections to establish where retrieval improves accuracy and where it does not.
Meet our team
Researcher
Muhammad Talal Majeed
Hi! I'm Talal, a tech student in Islamabad interested in closing the gap between the software industry and academia. I started with game development in 2018, moved through web development, and now work part-time in MLOps and DevOps. In 2026 I founded ecello.net, a place for collaborative research. My research sits in deep learning for document intelligence.
Researcher
Momena Akhtar
Hi! I'm Momena, a final-year Computer Science student at NUST. I'm currently researching Document VQA at DFKI, and on the product side I build and ship web and cloud applications. I like working across both: research questions that come out of real systems, and systems that get better because of what that research turns up. That mix is what keeps me interested.
Our blog
Coming soon
We are still writing up our first results. Papers and notes will be listed here as they go out.
Get in touch
If you are working on document understanding, retrieval, or vision-language models and want to compare notes, collaborate, or reproduce something we published, write to us.