Backend Engineer · Python · Django · PostgreSQL · Docker
LLM Evaluation · golden sets · LLM-as-judge · eval harnesses
CS undergrad at Scaler School of Technology & BITS Pilani · Bangalore, India
I build backend systems that actually ship. As a solo freelancer I took a B2B e-commerce platform from zero to production in 6 weeks — Django REST API, React frontend, PostgreSQL — now live at pronounjeans.com and serving 50+ active B2B users. Most of my work sits at the intersection of clean API design, concurrency correctness, and containerized deployment.
I also build evaluation for LLM systems — hand-labelled golden sets, blind LLM-as-judge scoring, and judge-vs-human agreement studies, because a system you can't measure isn't a system you can trust.
Currently: building CodeGraph — a static-analysis engine that maps code relationships using Python's AST, with Celery + Redis handling heavy parsing off the request thread.
Open to: backend / full-stack internships at early-stage startups where I can own features end-to-end, and contract work on AI evaluation and benchmark design.
Freelance · Live in production · Mar–Apr 2026
Designed, built, and deployed the entire platform solo, end-to-end.
- 50+ active B2B users on the live platform today
- 4 modular DRF apps — product catalog, cart, checkout, order management — exposing REST APIs to a React frontend
- 3-service deployment pipeline owned end-to-end: Django on Render Pro, React on Vercel, PostgreSQL on Supabase
- Diagnosed and fixed a production static-file outage via Django's
collectstaticpipeline + Whitenoise — eliminated 100% of frontend asset errors
Django Django REST Framework PostgreSQL Supabase React Whitenoise Render Vercel
Sept 2026
An AI support agent built on the Customer Support on Twitter dataset — and, more to the point, the evidence for whether it can be trusted.
- 10-intent taxonomy derived from the data, not assumed up front
- Hand-labelled golden set with a written codebook, plus human reply scores
- Blind LLM-as-judge harness, validated against human raters for agreement
- Documented failure analysis — including that the agent fabricated employee initials in 91% of replies — and a written account of why the headline metric misleads
Python LLM-as-judge golden sets inter-rater agreement failure analysis Make
In progress · June 2026 – Present
Code-parsing system built on Python's ast module to extract structural relationships from source files.
- Async parsing via Celery + Redis — resource-intensive jobs never block the main application thread
- Multi-container stack — frontend, backend, and Celery workers containerized and orchestrated with Docker Compose
- RESTful API serves parsed AST data and code metrics to a React + Vite frontend that renders the graphs
Python Celery Redis Docker Docker Compose React Vite
Dec 2025 – Feb 2026
Full-stack banking application — designed and built the React frontend and Django REST backend independently.
- Atomic fund transfers using Django ORM's
select_for_update()— prevents double-spend under concurrent requests - PostgreSQL schemas with integrity constraints enforcing valid financial states at the database level
- Structured into modular Django apps for clean separation of concerns
Django PostgreSQL REST APIs Docker React
Languages Python JavaScript SQL Java
Backend Django Django REST Framework Celery
Frontend React Vite
ML & Evaluation pandas Jupyter DuckDB golden-set design LLM-as-judge inter-rater agreement
Data & Infra PostgreSQL Redis Supabase Docker Docker Compose Linux Git Postman
Focus areas REST API design · async task processing · concurrency control · database schema design · containerization · evaluation design
Reach me at talindaga692@gmail.com — I reply fast.
