Skip to content

About

I am a Computer Vision Research Engineer at Roll.ai (San Francisco), working remotely. I build monocular systems for 3D reconstruction, depth and motion estimation, and controllable camera motion.

My work lives where perception meets geometry, recovering 3D structure and motion from ordinary video. Earlier, I researched safe multi-robot navigation at Information Technology University and built a license-plate recognition system deployed across seven countries at Hazen.ai.

I am moving toward robotics: motion planning and autonomous systems, carrying the same perception and geometry off the screen and into the physical world.

Research Interests

  • Computer Vision & Perception
  • 3D Reconstruction & Multi-View Geometry
  • Monocular Depth & Motion Estimation
  • Spatial AI
  • Robotics Perception
  • Motion Planning & Autonomous Systems

News

  • Jul 2025Building agentic LLM systems at Turing: multi-turn tool use and function calling.
  • May 2025Joined Roll.ai as a Computer Vision Research Engineer (remote).
  • Apr 2025Started trend-aware multi-asset forecasting research at Rawa'a AI.
  • Jan 2025Shipped a multi-country license-plate recognition system at 98% accuracy at Hazen.ai, commissioned by SDAIA.
  • Sep 2024Awarded the NGIRI national research grant for scalable multi-robot navigation.
  • Jun 2024Graduated with a B.S. in Computer Science from Information Technology University.
  • Fall 2023Taught calculus tutorials for 65+ students as a teaching assistant at ITU.
  • 2023Authored 60+ computer-vision and robotics tutorials at Educative, reaching 1M+ learners.
  • 2021Built Clean Street, a Unity-based endless runner, during the Mindstorm Studios fellowship.

Selected Work

Cinematic 3D Reconstruction & Camera Motion

Roll.ai · 2025–Present

A modular monocular system that coordinates depth, optical flow, segmentation, and reconstruction to recover cinematic camera motion from ordinary video. Segmentation-aware processing suppresses artifacts, and motion-guided ControlNet diffusion enables controllable camera trajectories, yielding stable, high-quality 2K output with a 3× inference speedup.

3D ReconstructionDiffusion / ControlNetCamera Motion Controlroll.ai

Safe Navigation of Multi-Agent Robotic Systems

ITU · B.S. Thesis · 2023–2025

A safety-aware framework for collision-free coordination of 5–50 robots in dense space. A dual-controller design pairs trajectory planning with a fallback safety controller, and a precedence-based decentralized strategy resolves conflicts without a central coordinator. Benchmarked on smoothness, makespan, travel distance, and collision severity.

Motion PlanningMulti-Agent SystemsSafe Autonomy

License-Plate Recognition at National Scale

Hazen.ai · SDAIA · 2024–2025

A recognition system deployed across seven Middle Eastern countries under smog, night, blur, and distortion. A variational autoencoder for OCR, paired with an 84M-sample synthetic-data pipeline, reaches 98% real-time accuracy and generalizes across countries from a single model, with no retraining.

OCRSynthetic DataGenerative Models

Trend-Aware Multi-Asset Forecasting

Rawa'a AI · 2025

Deep forecasting across 35+ assets using TiDE and Temporal Fusion Transformers over a 600+ feature dataset fusing market, macroeconomic, and alternative signals. A proposed Trend-Aware Loss jointly optimizes prediction error and directional accuracy.

Time-SeriesTransformers

Agentic LLM Systems

Turing · 2025–2026

Multi-turn agent-completion workflows spanning tool use, function calling, and real-world reasoning. Modeled assistant behavior across email, calendar, and maps-style tools with correct API-style execution across multi-step tasks.

LLM AgentsTool UseFunction Calling

Poverty Detection from Satellite Imagery

Intelligent Machine Lab · 2023

Predicting socioeconomic indicators from Sentinel-2 imagery and DHS survey data. Trained CNNs against Vision Transformers over 15K+ images, reaching 76% accuracy with ViTs, alongside an interpretability analysis of poverty-linked feature representations.

Vision TransformersRemote Sensing

Experience

  • Computer Vision Research Engineer, Roll.ai

    San Francisco · Remote · May 2025–Present

  • Undergraduate Research Assistant, Information Technology University

    Lahore, Pakistan · Sep 2023–Apr 2025

  • Associate Computer Vision Engineer, Hazen.ai

    Lahore, PK / Saudi Arabia · Jul 2024–Jan 2025

Education

B.S. in Computer Science, Information Technology University of the Punjab (ITU)

2020–2024

Lahore, Pakistan · CGPA 3.41 / 4.0

Thesis: Safe Navigation of Multi-Agent Robotic Systems

Coursework: Deep Learning, Computer Vision, Artificial Intelligence, Cyber-Physical Systems

Honors & Awards

  • BISP – HEC Fully Funded Scholarship

    2020–2024

    Higher Education Commission of Pakistan

    Full funding across the entire undergraduate degree.

  • NGIRI Research Grant

    Sep 2024

    Ministry of IT & Telecom, Pakistan

    Competitive national grant supporting research on scalable multi-robot navigation.

Technical Skills

Languages

  • Python
  • C++
  • MATLAB

Computer Vision

  • 3D Reconstruction
  • Monocular Depth
  • Optical Flow
  • Segmentation
  • OCR

3D Vision

  • Camera Motion Estimation
  • Novel View Synthesis
  • Multi-View Geometry
  • Temporal Consistency

Video AI

  • Motion Tracking
  • Video Segmentation
  • Diffusion Video Models
  • ControlNet

Robotics

  • Multi-Robot Navigation
  • Motion Planning
  • Collision Avoidance
  • Safety-Aware Planning

Machine Learning

  • Deep Learning
  • Vision Transformers
  • Reinforcement Learning
  • Multi-Agent Systems

Frameworks

  • PyTorch
  • TensorFlow
  • OpenCV
  • HuggingFace

Tools

  • Docker
  • Git
  • AWS
  • Linux
  • Unity

Contact

Open to research & engineering opportunities

I'm always glad to talk about perception, 3D vision, and robotics, whether it's research, a role, or a hard problem worth solving. Email is the fastest way to reach me.