trace talha-mumtaz

Talha Mumtaz

Python backend engineer · Islamabad, Pakistan

I design and ship production services end to end, from data model and REST API through background processing, third-party integrations, and deployment.

Experience

  • NUST
  • IBS Digital
  • AIO
2022AprJulOct2023AprJulOct2024AprJulOct2025AprJulOct2026AprJulOct

Associate Python Engineer

AIO

Attributes

span.id
aio.associate-python-engineer
organization
AIO
role
Associate Python Engineer
location
Islamabad, Pakistan
period
Jan 2026 – now
duration
9m

Stack

  • FastAPI
  • asyncio
  • OpenAI
  • Anthropic
  • DynamoDB
  • Prometheus
  • Grafana
  • AWS ECS

Backend owner for two production LLM services and two internal platform tools, running as containers on AWS ECS across four environments.

Events (7)

  1. layered re-architecture

    Re-architected two production FastAPI services into a layered routes, pipelines, agents, and services structure, giving each layer one responsibility and making the model and database boundaries independently testable.

  2. event-loop unblocking

    Eliminated event-loop blocking across async request paths: moved synchronous database work onto worker threads, replaced sequential awaits with bounded concurrent execution, and capped parallel model calls with semaphores so a single large request could no longer exhaust provider rate limits.

  3. provider-agnostic llm client

    Designed a provider-agnostic async LLM client supporting OpenAI and Anthropic, with configuration-driven model fallback, schema validation of model output, and error propagation that surfaces the real cause at the API layer instead of a generic 500; moved prompts out of code into versioned templates so copy changes no longer require a deployment.

  4. observability

    Instrumented both services against an internal observability library: per-stage pipeline tracking, and per-call token and cost logging persisted to DynamoDB, with service metrics pushed through a Prometheus pushgateway. Wrote the integration to degrade to no-ops so a monitoring outage slows a request rather than failing it.

  5. grafana dashboards

    Built Grafana dashboards over those metrics covering pipeline runs, failure rates, model latency, and token usage, giving the team one view of how the services behaved in production instead of reading logs after the fact.

  6. load testing + audit

    Built concurrent and sequential load-test suites backed by recorded-response fixtures, isolating service throughput from model latency; used the results to audit the codebase, document 23 correctness and scalability defects, including success responses returned after failed database lookups and outbound calls with no timeout, and delivered the fixes.

  7. onboarding

    Onboarded junior engineers onto these codebases, writing onboarding documentation and walking them through the architecture and the async conventions the code depends on.

Child spans

Skills

languages
Python, SQL, Bash
backend
FastAPI, Pydantic, SQLAlchemy (sync + async), Alembic, Celery, asyncio, Uvicorn, REST API design
databases
PostgreSQL, pgvector, Redis, DynamoDB
ai.llm
AWS Bedrock, OpenAI, Anthropic, RAG, ReAct agents, LangChain, LangGraph, Semantic caching, Prompt engineering & versioning
cloud.infra
AWS (ECS, S3, Lambda, DynamoDB, CloudWatch, SES, Bedrock), Docker, Docker Compose, GitHub Actions, Linux, Git
observability
Prometheus, Grafana, CloudWatch, LangSmith, Langfuse, Structured logging, LLM token & cost tracking
testing
pytest, pytest-asyncio, Load testing
frontend
HTML, CSS, JavaScript, React (familiar)

Certifications

  • LangChain for LLM Application Development

    DeepLearning.AI

  • AI Agents in LangGraph

    DeepLearning.AI

  • Serverless Agentic Workflows with Amazon Bedrock

    DeepLearning.AI

  • Introduction to DevOps

    Coursera

Get in touch

Reach me by email or on LinkedIn.