Talha Mumtaz
Python backend engineer · Islamabad, Pakistan
I design and ship production services end to end, from data model and REST API through background processing, third-party integrations, and deployment.
Experience
- NUST
- IBS Digital
- AIO
Talha Mumtaz
Python backend engineer
Attributes
- span.id
- talha-mumtaz
- name
- Talha Mumtaz
- focus
- Python backend engineer
- location
- Islamabad, Pakistan
- roles
- 2
- projects
- 2
- period
- Nov 2021 – now
- duration
- 4y 11m
I design and ship production services end to end, from data model and REST API through background processing, third-party integrations, and deployment. Recent work centers on multi-tenant applications built on FastAPI and LLMs: retrieval pipelines, agent workflows, and Celery task processing, together with the architecture and concurrency work that keeps them predictable as load grows.
Spans in this trace
BE Software Engineering
National University of Sciences and Technology
Attributes
- span.id
- nust.be-software-engineering
- institution
- National University of Sciences and Technology
- degree
- BE Software Engineering
- result
- CGPA 3.32 / 4.00
- period
- Nov 2021 – May 2025
- duration
- 3y 7m
Backend & GenAI Engineer
IBS Digital
Attributes
- span.id
- ibs.backend-and-genai-engineer
- organization
- IBS Digital
- role
- Backend & GenAI Engineer
- location
- Islamabad, Pakistan
- period
- May 2025 – Jan 2026
- duration
- 9m
Stack
- FastAPI
- Celery
- Redis
- SQLAlchemy
- PostgreSQL
- Docker
Events (4)
multi-tenant genai platform
Built and maintained a multi-tenant GenAI backend platform on FastAPI, Celery, SQLAlchemy, and Docker, supporting conversational AI, document ingestion, analytics, and web-scraping workflows.
rag + agents
Engineered retrieval-augmented generation and agentic workflows for contextual data retrieval and automated decision-making across the platform.
celery offload
Moved long-running document processing and crawling onto Celery with Redis as the broker, keeping request paths responsive and making job status, retries, and failures visible to users.
tenant isolation
Implemented tenant-scoped data separation and asynchronous task execution for enterprise deployments.
Child spans
Multi-Tenant GenAI Support Assistant
Production support assistant across web, mobile, WhatsApp, and Teams
Attributes
- span.id
- ibs.support-assistant
- project
- Multi-Tenant GenAI Support Assistant
- built at
- IBS Digital
- period
- 2025
Stack
- FastAPI
- PostgreSQL + pgvector
- Redis
- Celery
- AWS Bedrock
- LangChain
- LangGraph
A production support assistant serving multiple tenants across web, mobile, WhatsApp, and Microsoft Teams. Owned the GenAI and backend layers.
Events (6)
routing layer
Built the routing layer: one classification pass per message that applies content guardrails, condenses the conversation into a standalone query, and detects language, sentiment, and product entities, then dispatches to the right handler, whether knowledge-base retrieval, a LangGraph agent, or a scheduling and human-handoff flow, instead of sending every message down one expensive path.
react agent
Implemented the ReAct agent in LangGraph with a Redis checkpointer for per-session conversation state, per-tenant compiled-graph caching, and a fallback node that returns tool errors to the model to retry rather than failing the turn.
semantic cache
Added a semantic cache on Redis using an HNSW vector index with cosine similarity, keyed on the condensed query, so repeat questions skip retrieval and generation entirely; the embedding call runs off the event loop and any cache failure falls back to full retrieval.
ingestion
Built document and website ingestion as Celery tasks over Redis, extracting PDF, DOCX, and text content from S3 and crawled pages, then chunking and embedding it into pgvector with per-job status tracking, so long ingestions never block the API.
model controller
Abstracted model access behind a controller covering Bedrock, SageMaker, and Azure OpenAI, and scored answers for relevance and groundedness per message to track answer quality over time.
outbound email
Implemented the outbound email flows for sales handoff, support escalation, and negative-sentiment alerts, rendering the conversation transcript into the notification.
Associate Python Engineer
AIO
Attributes
- span.id
- aio.associate-python-engineer
- organization
- AIO
- role
- Associate Python Engineer
- location
- Islamabad, Pakistan
- period
- Jan 2026 – now
- duration
- 9m
Stack
- FastAPI
- asyncio
- OpenAI
- Anthropic
- DynamoDB
- Prometheus
- Grafana
- AWS ECS
Backend owner for two production LLM services and two internal platform tools, running as containers on AWS ECS across four environments.
Events (7)
layered re-architecture
Re-architected two production FastAPI services into a layered routes, pipelines, agents, and services structure, giving each layer one responsibility and making the model and database boundaries independently testable.
event-loop unblocking
Eliminated event-loop blocking across async request paths: moved synchronous database work onto worker threads, replaced sequential awaits with bounded concurrent execution, and capped parallel model calls with semaphores so a single large request could no longer exhaust provider rate limits.
provider-agnostic llm client
Designed a provider-agnostic async LLM client supporting OpenAI and Anthropic, with configuration-driven model fallback, schema validation of model output, and error propagation that surfaces the real cause at the API layer instead of a generic 500; moved prompts out of code into versioned templates so copy changes no longer require a deployment.
observability
Instrumented both services against an internal observability library: per-stage pipeline tracking, and per-call token and cost logging persisted to DynamoDB, with service metrics pushed through a Prometheus pushgateway. Wrote the integration to degrade to no-ops so a monitoring outage slows a request rather than failing it.
grafana dashboards
Built Grafana dashboards over those metrics covering pipeline runs, failure rates, model latency, and token usage, giving the team one view of how the services behaved in production instead of reading logs after the fact.
load testing + audit
Built concurrent and sequential load-test suites backed by recorded-response fixtures, isolating service throughput from model latency; used the results to audit the codebase, document 23 correctness and scalability defects, including success responses returned after failed database lookups and outbound calls with no timeout, and delivered the fixes.
onboarding
Onboarded junior engineers onto these codebases, writing onboarding documentation and walking them through the architecture and the async conventions the code depends on.
Child spans
ProgressBoard
Team delivery & reporting platform
Attributes
- span.id
- aio.progressboard
- project
- ProgressBoard
- built at
- AIO
- period
- 2026
Stack
- FastAPI
- PostgreSQL
- DynamoDB
- S3
- Jira & Notion APIs
- React SPA
A delivery tracker for the engineering department: sprint boards, per-person progress, and automated weekly status reports. Designed and built the backend end to end.
Events (4)
hierarchical access
Designed the hierarchical access model, an org tree of arbitrary depth where a user sees everyone below them and never peers or managers, resolved in a single recursive CTE rather than repeated queries.
authentication
Implemented authentication with Argon2id password hashing, short-lived JWT access tokens, and rotating refresh tokens stored hashed and delivered in an httpOnly cookie, with constant-time login so a wrong email cannot be distinguished from a wrong password.
jira + notion sync
Integrated Jira Cloud and Notion behind async clients that paginate, back off on rate limits, and pin the API version, syncing issues, sprints, and the tasks tracker into a normalized Postgres model on a scheduled job.
llm status reports
Generated weekly status reports with an LLM over pre-aggregated facts rather than raw task dumps, keeping token cost flat as the board grows; reports render to HTML and PDF, go out by email, and archive to S3.
Skills
- languages
- Python, SQL, Bash
- backend
- FastAPI, Pydantic, SQLAlchemy (sync + async), Alembic, Celery, asyncio, Uvicorn, REST API design
- databases
- PostgreSQL, pgvector, Redis, DynamoDB
- ai.llm
- AWS Bedrock, OpenAI, Anthropic, RAG, ReAct agents, LangChain, LangGraph, Semantic caching, Prompt engineering & versioning
- cloud.infra
- AWS (ECS, S3, Lambda, DynamoDB, CloudWatch, SES, Bedrock), Docker, Docker Compose, GitHub Actions, Linux, Git
- observability
- Prometheus, Grafana, CloudWatch, LangSmith, Langfuse, Structured logging, LLM token & cost tracking
- testing
- pytest, pytest-asyncio, Load testing
- frontend
- HTML, CSS, JavaScript, React (familiar)
Certifications
LangChain for LLM Application Development
DeepLearning.AI
AI Agents in LangGraph
DeepLearning.AI
Serverless Agentic Workflows with Amazon Bedrock
DeepLearning.AI
Introduction to DevOps
Coursera
Get in touch
Reach me by email or on LinkedIn.