ChatWith.study
AI-powered learning and research companion
An AI-powered study and research platform that helps students and researchers discover project topics, understand course materials, query uploaded documents, and work with academic content through contextual AI.
01 / CONTEXT
The problem
Students and researchers often spend substantial time searching for appropriate final-year project topics and extracting useful information from large collections of academic material. Existing tools may provide generic answers without sufficient grounding in the user's actual course or research material.
02 / APPROACH
How I built it
Uploaded documents are processed into semantically meaningful chunks, embedded into pgvector, and retrieved through hybrid search. Retrieved content is reranked before being provided to the language model, while citation verification is enforced at the generation layer.
03 / CONSTRAINTS
What had to hold
Large academic documents 300+ page textbook PDFs Tables and mathematical content Unstructured document formatting Hallucination prevention Citation accuracy
04 / OBJECTIVES
Definition of done
• Provide contextual answers grounded in uploaded material • Generate useful study resources • Help students discover research topics • Support research-document interaction • Provide page-level citations • Maintain conversational context
05 / THE SYSTEM
Architecture
Document ingestion pipeline Uploaded documents are parsed, segmented, and prepared for retrieval. Hybrid retrieval Vector similarity and keyword BM25 search are combined to improve retrieval coverage. Semantic chunking Document boundaries and heading context are preserved to improve retrieval quality. Reranking Candidate results are reranked before entering the model context. Citation-grounded generation The conversational AI is constrained to retrieved material and produces page-level citations. Streaming chat The client receives conversational responses through SSE while retrieval tools execute server-side.
06 / DECISIONS
Key decisions
Use hybrid search rather than vector search alone Why: Academic documents contain terminology where keyword matching can complement semantic similarity. Use semantic chunking Why: Preserving conceptual and heading boundaries improves retrieval quality. Rerank retrieved results Why: Top-k retrieval quality must be improved before information reaches the language model. Require citation verification Why: Academic assistance requires traceable answers rather than unsupported model generations.
07 / SHIPPED
Deliverables
• Document upload (in-progress) — Core capability of the product (see features). • PDF parsing (in-progress) — Core capability of the product (see features). • PPTX parsing (in-progress) — Core capability of the product (see features). • DOCX parsing (in-progress) — Core capability of the product (see features). • Lecture-note analysis (in-progress) — Core capability of the product (see features).
08 / LESSONS
What I'd keep
• Chunk size and semantic boundary detection can matter more to retrieval quality than raw model capability. • Academic AI systems need grounding and citation mechanisms rather than relying solely on model intelligence. • Large documents require retrieval architecture designed around structure, not just token volume.
What changed
Sub-second response time
Semantic search response
to date
300+ pages
Page documents handled
to date
// START A PROJECT
Building something with similar constraints?
If it's in your critical path and you'd rather not learn its failure modes live, let's talk about it.