AI / ML Engineer / Islamabad, Pakistan

Shujaan Azhar

Lesson plans for Pakistani classrooms. A call-centre voice agent that runs on one GPU and costs nothing to operate. A tax optimiser that never sees your salary. I build the AI, and the harness that proves it works.

  • 14 checks an exam clears before it ships Core AI
  • 0.3 s speech round-trip, on one local GPU Voice
  • ×3 negotiations run at once, not in turn Agentic
  • 0 bytes of your salary that leave the browser Development
Currently building education AI at Taleemabad
scroll
§ 01

Work

9 built out · 11 more

Four things I actually do, each one demonstrated below rather than described.

  1. 01 Development Open source 2026 Solo

    TaxBacha

    A tax optimiser for Pakistani salaried filers — not “what do I owe”, but “what could I owe instead”.

    Befiler and FBR IRIS optimise for filing compliance. TaxBacha optimises the bill: it computes your liability against the FY 2025-26 salaried slabs, then ranks the legal instruments that reduce it — VPS under §63, Zakat under §60, approved donations under §61 — by rupees saved. No API keys, no account, no salary leaving the machine.

    Slabs
    FY 2025-26
    Keys required
    none
    Data leaves device
    never
    • Next.js 16
    • React 19
    • TypeScript
    • Tailwind 4
    • shadcn/ui
    Read the source
    tax-engine.ts FY 2025-26 · salaried

    six brackets · 9% surcharge above PKR 10M

    Tax as filed 861,000
    Optimised
    Tax saved
    bracket reached unreached Zakat deducts from income; the rest are credits at your average rate.

    This is the real engine. The slab function and every credit rule below are the exact logic from app/lib/tax-engine.ts, ported line-for-line. Drag the salary.

  2. 02 Voice & Realtime Open source 2026 Solo

    VICIdial AI Voice Agent

    A call-centre voice agent with barge-in that runs entirely on my own GPU. No cloud, no carrier, no PSTN.

    Speech recognition, the language model, and speech synthesis all run locally; the “customers” are SIP softphones on a private subnet. It hooks into Asterisk audio over AudioSocket — an 8 kHz PCM TCP stream — and VICIdial’s campaign, lead and reporting layer wires in on top. Total running cost is zero and no audio ever leaves the machine.

    Response latency
    ~0.3 s
    Running cost
    $0
    Built on
    RTX 4060 · 8 GB
    • faster-whisper
    • Ollama · qwen2.5:7b
    • Piper TTS
    • Asterisk 18
    • KVM/libvirt
    • Python 3.12
    Read the source
    deployment topology zero PSTN
    Host 192.168.122.1
    • VAD energy · adaptive floor CPU
    • STT faster-whisper small.en GPU
    • LLM Ollama qwen2.5:7b GPU
    • TTS Piper lessac-medium CPU
    KVM guest 192.168.122.10
    • Asterisk 18 SIP / RTP · AudioSocket
    • VICIdial campaigns · dialer daemons
    • MySQL + Perl lead database
    • Apache / PHP agent web UI

    libvirt NAT bridge virbr0 · 192.168.122.0/24 · no carrier, no DID, no cloud API

    GPU CPU round trip ~0.3 s with barge-in cost $0

    The actual deployment topology — host and guest on one libvirt NAT bridge, no PSTN anywhere in it. The pulse is a turn of the audio loop.

  3. 03 Voice & Realtime Open source 2026 Solo

    Jarvis — clap-activated assistant

    Double-clap; it wakes, listens, answers out loud, and goes back to sleep. Every hosted hop has a local fallback.

    Clap detection runs continuously on MFCC embeddings matched against a voice-print you calibrate yourself, so it fires on your clap and not a door closing — without holding the mic open for transcription. It keeps a persistent memory of you in a plain-Markdown Obsidian vault, reads your calendar, and hands code changes to Claude Code, routed to the right repo by spoken alias.

    Wake word
    your own clap
    Local fallbacks
    STT + TTS
    Memory
    Markdown vault
    • Python 3.12
    • MFCC / librosa
    • Groq
    • faster-whisper
    • Piper
    • Claude Code
    Read the source
    degradation path
    1. Wake MFCC voice-print always on device
    2. Hear Groq Whisper faster-whisper falls back
    3. Think Groq LLM · Claude needs network
    4. Speak Groq Orpheus Piper falls back
    5. Remember Markdown vault always on device

    Clap detection runs continuously against a voice-print you calibrate yourself — it fires on your clap, not a door.

    active path standby press network up to pull the plug

    Hosted path on top, local fallback underneath. Toggle the network to watch it degrade instead of die.

  4. 04 Voice & Realtime Open source 2026 Solo

    Decepticon — autonomous meeting agent

    Joins your Google Meet calls from your calendar, listens, decides whether to speak, and replies in your cloned voice.

    The hard part is not speaking — it is knowing when not to. Transcription runs locally on GPU with Faster-Whisper, a local Mistral decides whether the moment warrants a reply, and Chatterbox synthesises the answer with zero-shot cloning from a ten-second reference clip. Audio moves over virtual PulseAudio devices, so there is no physical mic or speaker in the loop.

    Runs
    fully local, on-GPU
    Voice clone
    zero-shot, ~10 s ref
    Hardware needed
    none
    • Faster-Whisper
    • Ollama · Mistral
    • Chatterbox TTS
    • PulseAudio
    • CUDA
    Read the source
    turn decision sample meeting
    1. Sara so who is picking up the migration this sprint hold
    2. Omar I can take the schema half of it hold
    3. Sara right — and the rollback plan? hold
    4. Omar Shujaan wrote that one last quarter hold
    5. Sara Shujaan, can you walk us through it? speak
    6. agent speaking — cloned voice tts
    turns heard 0 turns spoken 0 restraint is the feature
    hold speak Faster-Whisper → Mistral → Chatterbox, all on-GPU

    The decision the agent actually makes on every turn boundary. Most of them are “stay silent”.

  5. 05 Agentic Systems Open source 2026 Solo · hackathon

    Sahulat — service orchestrator

    Seven agents that replace the thirty minutes of phone calls it takes to book a plumber in Pakistan.

    Type a request in Urdu, Roman Urdu, or English. Intake structures it, Discovery filters the directory by haversine distance in pure Python, Ranking scores candidates, and then the Coordinator spawns three concurrent negotiations — each provider played by an in-character LLM persona — asking can you come, what time, what price. It books the best committed offer. Built for the AI Seekho 2026 Antigravity Hackathon.

    Agents
    7
    Negotiations
    3, concurrent
    End to end
    ~5–8 s
    • FastAPI
    • Llama 3.3 70B
    • Llama 3.1 8B
    • Groq
    • Expo / React Native
    • WhatsApp Cloud API
    Read the source
    coordinator agent sample run · 3 × Llama 3.3 70B
    rank score
    • Ahmed R. electrician · 1.2 km · 4.6★
      PKR ETA
    • Bilal K. electrician · 2.8 km · 4.8★
      PKR ETA
    • Kashif M. electrician · 0.9 km · 4.2★
      PKR ETA
    coordinator waiting on all three…
    in flight committed 7 agents · end to end ~5–8 s

    The Coordinator’s three parallel negotiations, and the scoring rule that ranked the providers going in.

  6. 06 Agentic Systems Deployed 2026 Solo · personal

    Sheikhspeare — Slack agent

    An always-on Slack agent in iambic pentameter that knows who is asking, and refuses the ones who shouldn’t.

    The interesting problem in a personal assistant that lives in a shared workspace is not personality — it is containment. Identity is resolved per message before any tool is reachable, so the private toolset and the owner’s context are gated at the boundary rather than guarded by a prompt. It runs 24/7 as a Socket Mode worker with an owner-only conversation log.

    Uptime
    24/7 worker
    Gate
    per-message identity
    Voice
    Shakespearean
    • Claude Agent SDK
    • Python
    • Slack Socket Mode
    • Railway
    • Docker
    identity gate per message

    both send “what did I say I'd ship this week?”

    @shujaan owner
    resolve
    owner
    toolset
    3 mounted
    • owner context mounted
    • private memory mounted
    • conversation log mounted
    Answers — in iambic pentameter, with the receipts.
    @anyone-else guest
    resolve
    not owner
    toolset
    0 mounted
    • owner context withheld
    • private memory withheld
    • conversation log withheld
    Declines — courteously, and in verse. It will not say whose context it holds.

    The gate runs before the model is handed a toolset — refusal is a property of the boundary, not a line in a prompt asking it nicely.

    The same question, asked by two people. Identity resolves before the toolset is even visible.

  7. 07 Core AI Live in production 2025— Design & build · core engineer

    AI Lesson Plan Generator

    Curriculum in, classroom-ready bilingual lesson plans out — with a reviewer that can send them back.

    Multi-grade, English and Urdu, with auto-generated illustrations. The part that matters is the gate: an AI reviewer scores every plan against a pedagogy rubric before it ships, and a failing plan is regenerated rather than delivered. The async pipeline is migrating to a LangGraph state machine so each stage is inspectable and retryable on its own.

    Languages
    English · Urdu
    Quality
    gated, not assumed
    Every call
    traced
    • Python
    • FastAPI
    • LangGraph
    • Redis
    • Langfuse
    • AWS EC2
    generation pipeline LangGraph state machine
    1. intake curriculum + grade
    2. draft plan, EN + UR
    3. figures illustrations
    4. review pedagogy rubric
    5. deliver to the classroom
    rubric
    • objectives
    • Bloom's coverage
    • standards alignment
    • assessment for learning
    • bilingual parity

    idle

    in flight passed the gate a failing plan is regenerated, not shipped

    The generate → review → gate loop. A plan below threshold does not ship; it goes back around.

  8. 08 Core AI Live in production 2025— Design & build · core engineer

    AI Exam Generator

    Any teaching content becomes a standards-aligned exam — as structured data other platforms consume over HTTP.

    Multiple question types, bilingual, with figures. Because it returns structured exam data rather than a rendered PDF, it works as a reusable content API instead of a one-off document maker. A fourteen-check, standard-mapped reviewer grades every exam on a calibrated scale before it is returned.

    Reviewer checks
    14
    Returns
    structured data
    Consumed by
    other platforms
    • Python
    • FastAPI
    • Redis
    • Langfuse
    • AWS EC2
    reviewer harness 14 standard-mapped checks
    exam under review
    • multiple question types
    • English + Urdu
    • figures
    • standards-mapped

    waiting for an exam…

    checking passed graded on a calibrated scale before the payload is returned

    The reviewer’s check surface. Every generated exam is graded against all fourteen before it leaves the service.

  9. 09 Core AI Shipped 2026 ML engineer

    Vision Exam Parsing

    Swapped a GPT-4o vision call out of a production parser for a YOLO detector I fine-tuned myself.

    Handwritten exam pages had to be segmented and cropped at scale, and a hosted vision model was the wrong tool for a fixed, repetitive layout problem. I curated a domain dataset, fine-tuned YOLO v11/v12 over 100+ epochs, and paired it with OCR — self-hosted, so throughput is a function of my own hardware rather than someone’s rate limit. It ships behind FastAPI with APM tracing and process-pooled workers.

    Training
    100+ epochs
    Replaced
    a hosted vision call
    Now runs
    self-hosted
    • YOLOv12
    • Ultralytics
    • PyTorch
    • Google Vision OCR
    • FastAPI
    • Elastic APM
    YOLOv12 · fine-tuned self-hosted
    marks box question diagram answer · handwritten
    regions cropped → OCR
    • marks box 46 × 24
    • question 240 × 54
    • diagram 100 × 70
    • answer 240 × 70

    GPT-4o vision call detector I trained

    100+ epochs on a curated domain dataset. Throughput is now a function of my hardware, not someone's rate limit.

    What the detector does to a page: regions found, classified, and cropped for OCR downstream.

Also shipped

  • Core AI

    BenchLP-PKResearch paper · in progress

    An IEEE-format benchmark of LLMs for Pakistani primary lesson-plan generation (Grades 1–5, four subjects), proposing a rubric grounded in Bloom’s Taxonomy and Assessment for Learning — and evaluating models as reviewers, not just generators.

    private
  • Core AI

    Self-improving prompt & eval loopR&D

    Calibrate an AI reviewer against human-expert annotations until it agrees with them, then let that agreement score drive automated prompt optimisation — versioned registry and score store underneath, human promotion gate on top.

    private
  • Core AI

    Testing Urdu LLMsEvaluation study

    Benchmarking Gemma, Qwen and Llama against commercial models for Urdu instruction-following — a low-resource language where the leaderboard rankings stop predicting anything.

    private
  • Core AI

    On-device offline lesson plansEdge AI

    Generation that runs entirely on the phone for classrooms with no connectivity — tiered Gemma models in React Native, sized against a real device fleet rather than a flagship.

    private
  • Core AI

    Ask-My-TextbooksAgentic RAG

    A retrieval agent that answers and writes reports over textbooks — shipped with its own evaluation harness and test set, because a RAG system without one is a demo.

    source
  • Core AI

    Evaluation platformEval infrastructure

    Comparative grading across GPT-4, Gemini and Claude, with the Claude Batch API doing the bulk work at a fraction of the interactive cost, surfaced through Streamlit dashboards.

    private
  • Core AI

    Urdu–English TranslatorSeq2Seq · coursework

    Sequence-to-sequence translation with Bahdanau attention in PyTorch, from scratch — 81 GLEU on 15k+ sentence pairs against a baseline transformer.

    private
  • Agentic Systems

    Sketchless · FloorForgeGenerative 3D · WIP

    A floor-plan to furnished-3D-scene pipeline built Claude-Code-native — agent skills and MCP servers driving Unreal Engine 5 and SketchUp, with a fixed benchmark scene as the regression test.

    private
  • Voice & Realtime

    TeravoxVoice bot

    An insurance call bot — Ultravox for transcription, Qwen3-4B for intent and response, orchestrated as a LangGraph workflow.

    private
  • Voice & Realtime

    VocalCraftMulti-modal pipeline

    Whisper, LLaMA and a LoRA-fine-tuned SDXL chained into a pamphlet generator — cut design time 80%, tested on 100+ menu items, deployed on AWS.

    private
  • Development

    PSX Dividend CalculatorWeb app

    Tracks Pakistan Stock Exchange dividends against live data, with withholding tax handled correctly by filer status — the detail every other tracker gets wrong.

    source
§ 02

Writing

papers, not just projects
  • Paper

    Beyond Generation: Evaluating LLMs as Pedagogical Generators and Reviewers for Primary Lesson Plans in a Multilingual South Asian Context

    Sole-authored preprint — targeting IEEE FIE 2026, Limassol, Cyprus · 2025

    A validated pedagogical rubric (Grades 1–5; English, Science, Maths, Urdu) and a benchmark of seven open-weight LLMs as both generators and reviewers — a generator–reviewer refinement loop lifted every model above 88% of the rubric maximum.

  • Blog

    A case study on optimizing LLMs for Yemeni Arabic

    Taleemabad Blog · 2026

    How a generation–review pipeline got tuned for Yemeni Arabic — frontier models down to self-hosted small LMs, the prompt change that fixed script rendering, and a dual-judge evaluation loop underneath it.

    drafting
§ 03

About

Making a model produce something is the easy half. Building the thing that can tell you it’s wrong — that’s the job.

I work on education AI at Taleemabad, where the lesson-plan and exam generators I build are used by teachers and gated by reviewers I also built — a plan that fails its rubric is regenerated, not delivered. Outside that, I keep pulling speech pipelines apart: a VICIdial voice agent with sub-second latency and no cloud dependency, a meeting bot that mostly decides to stay quiet, an assistant that wakes on the sound of my own clap. The through-line is that I would rather ship something measurable than something impressive.

Domains
Core AI · agentic · voice · development
Research
BenchLP-PK — benchmarking LLMs for Pakistani lesson plans
Based in
Islamabad, Pakistan
Currently
AI/ML Engineer, Taleemabad
§ 04

Experience

and where the time went
  1. Taleemabad

    Current

    AI Engineer · Islamabad

    Jan 2025 – Present

    • Architected and independently own a closed-loop agentic pipeline — generation → AI review → annotation → assessment — serving 1,000+ lesson plans/day and 400+ exams/week across 370+ schools. Own ~94% of the lesson-plan codebase and 100% of the exam-generator codebase, by git blame.
    • Designed the standards-mapped review rubrics that lifted AI-output quality from 50–60% to 88–95% — charted below.
    • Benchmarked 6+ model families, including self-hosted open-weight models on a dedicated GPU box — one now outperforms production GPT-4o on a low-resource-language task.
    review quality — before the rubric gate → after Taleemabad, 2025
    50–60%before — free-form grading
    88–95%now — rubric-gated
    −60% context usage, after the graph-based harness
    +80% task-completion accuracy, same harness
    −22% inference cost, via prompt caching
    1,000+ lesson plans generated / day
    400+ exams generated / week
    370+ schools served

    Same generator, same model family — the only variable is whether output is scored against an auditable rubric before it ships.

  2. Funavry Technologies

    AI Intern · Islamabad

    Jul 2024 – Aug 2024

    • Designed prompt-engineering workflows that lifted generative output quality 30%.
    • Built training pipelines over 100k+ Solidity smart-contract samples (Hugging Face, PyTorch) with RAG for code analysis.
  3. Xaleta Technologies

    Associate Data Scientist

    Apr 2024 – Jun 2024

    • Improved GPT-3.5/4 responses via RLHF scoring across 1,500+ outputs — consistency up 20%.
  4. AIM Research Lab — FAST University

    AI Intern

    Jun 2023 – Sep 2023

    • Built an AI cricket-umpire system — 96% accuracy across 200+ match videos, using CNNs and temporal analysis.

National University of Computer and Emerging Sciences (FAST)

BS, Computer Science · Islamabad

Aug 2021 – Jun 2025

Relevant coursework
  • Generative AI
  • Digital Image Processing
  • NLP
  • MLOps
Certifications
  • Machine Learning SpecializationStanford Online · 2023
  • NLP SpecializationDeepLearning.AI · 2023
  • Build Basic GANsDeepLearning.AI · 2024
§ 05

How I work

with receipts
01

The gate is the product

Anyone can generate a lesson plan. The work is the reviewer that scores it against a rubric and sends it back.

A plan below threshold is regenerated, not delivered · 14 checks before an exam is returned

02

Own the hardware path

A hosted API is a dependency with a price and an outage. Where a model can run on my own GPU, it does.

Voice agent at ~0.3 s round-trip for $0 · a fine-tuned YOLO replacing a GPT-4o vision call

03

Degrade, don’t die

Systems meet bad networks and rate limits. The interesting design question is what still works when they do.

Jarvis falls back to faster-whisper and Piper locally · Sahulat has a regex path when the LLM is down

04

Refuse at the boundary

A prompt asking a model to keep a secret is not access control. Resolve identity before the tools are reachable.

Sheikhspeare mounts its private toolset per message, after the sender is known

§ 06

Stack

and where each of it went
Speech
  • faster-whisper
  • Whisper
  • Piper
  • Chatterbox
  • Orpheus
  • Ultravox
  • VICIdial agent
  • Jarvis
  • Decepticon
  • Teravox
LLM serving
  • Ollama
  • Groq
  • OpenRouter
  • Claude
  • Gemini
  • GPT-4o
  • Qwen
  • Llama
  • Gemma
  • Mistral
  • every project on this page
Agents
  • Claude Agent SDK
  • LangGraph
  • LangChain
  • MCP
  • Sheikhspeare
  • Sahulat
  • LP pipeline
  • FloorForge
Evaluation
  • Langfuse
  • DSPy
  • rubric design
  • human calibration
  • Claude Batch API
  • Exam reviewer
  • BenchLP-PK
  • self-improving loop
Vision
  • YOLOv11 / v12
  • Ultralytics
  • PyTorch
  • Google Vision OCR
  • Exam parsing
  • CV coursework
Backend
  • Python
  • FastAPI
  • Redis
  • Postgres / Supabase
  • BigQuery
  • LP + exam generators
  • Sahulat
Frontend
  • TypeScript
  • Next.js 16
  • React 19
  • Tailwind 4
  • shadcn/ui
  • React Native
  • Astro
  • TaxBacha
  • on-device LP app
  • this site
Infrastructure
  • Docker
  • AWS EC2 / S3
  • GitHub Actions
  • Railway
  • Vercel
  • KVM / libvirt
  • Asterisk
  • VICIdial agent
  • Sheikhspeare
  • production services
Observability
  • Grafana
  • Prometheus
  • ELK
  • Elastic APM
  • Langfuse
  • LP + exam generators
  • production services
§ 07 — Contact

Let’s build something
worth measuring.

Open to collaborations and conversations about applied AI, evaluation, and education technology.