Hi, I'm Daud Shah

I am a

Computer Vision & AI engineer building multimodal medical AI, object detection, Object segmentation, object tracking, OCR's, VLMs, and research manuscripts ready for publication. Recent industry experience at CCRIPT Agency and Neuralogic(USA). Based in Islamabad.

Daud Shah portrait
CV / AI
PyTorch
YOLO
OpenCV
0+
Years experience
0
Companies
0+
Projects built
0
Freelance platforms

Experience

Experience & education

Six months of computer vision engineering at CCRIPT Agency and Neuralogic (Jan–Jun 2026); research internship at NCAI (completed); degree work in parallel.

Computer Vision Engineer
CCRIPT Agency
Islamabad, Pakistan
Jan 2026 – Jun 2026
  • End-to-end object detection and segmentation on client imagery using YOLO-family models — from dataset design through training, validation, and delivery-ready inference.
  • Structured data annotation and QA workflows in CVAT and Roboflow, keeping labels consistent from raw images to model-ready datasets.
  • Vision-language model (VLM) experiments for richer scene understanding and assisted review on complex visual domains.
  • Streamlit web apps so clients and stakeholders can upload images, run models, and inspect outputs without engineering overhead.
  • Computer vision on architectural drawings and floor plans — detecting layout elements, symbols, and regions for inspection and automation workflows.
  • Collaborating with the team to turn prototypes into maintainable, client-facing deliverables.
Computer Vision Engineer
Neuralogic
Texas, USA (remote)
Jan 2026 – Jun 2026
  • Object detection and model benchmarking (YOLO family, PyTorch) on US-facing computer vision initiatives in a remote, distributed team.
  • Dataset curation and annotation standards across CVAT and Roboflow — quality checks that keep training data reliable at scale.
  • Exploring VLMs for technical and document-like imagery where classical detectors benefit from semantic context.
  • Interactive Streamlit demos for rapid proof-of-concept, model comparison, and stakeholder feedback before full integration.
  • CV pipelines for architectural and technical drawings — layout parsing, element detection, and structured outputs for downstream use.
  • Reproducible train/eval workflows, clear documentation, and async coordination across time zones.
AI & Computer Vision Engineer — Intern
NCAI Lab, UET Peshawar
Peshawar, Pakistan
Nov 2024 – Jan 2026
  • Traffic and safety CV: ambulances, rickshaws, license plates (YOLOv8 / v11).
  • End-to-end annotation → training → evaluation; Roboflow experiment flows.
  • Video understanding (MMAction2) and OCR (PaddleOCR) with the research group.
Bachelor of Computer Science
University of Agriculture, Peshawar
Peshawar, Pakistan
Nov 2022 – Nov 2026 (expected)
Final year · CGPA 3.57 — coursework and projects aligned with AI/Computer Vision/NLP Computer Science, statistics, and data-driven systems.
FSc Pre-Engineering
Quaid-e-Azam Group of Schools & Colleges
Pakistan
2020 – 2022
Grade B — foundation for engineering and CS track.

Research & publications

Manuscripts and systems ready for journal submission — multimodal medical AI, efficient video recognition, and clinical imaging. Lab work at NCAI complemented this with traffic CV and deployable pipelines.

IEEE JBHI · In preparation

Multi-Modal Chest X-Ray Classification

Final Year Project · University of Agriculture, Peshawar · 2025

ViT + Bio-ClinicalBERT with four-layer bidirectional cross-attention on 172,202 image–report pairs from two hospitals. 98.52% macro-AUC across 15 pathology labels; cross-hospital gap 1.30%; Grad-CAM interpretability.

IEEE Access · Manuscript prepared

DAUD-NET v2 — Efficient Video Action Recognition

Dual-attention transformer · UCF-101

Sparse local–global temporal attention with adaptive temporal pooling. 97.62% accuracy on UCF-101 at 71 GFLOPs — 44% fewer attention computations than full attention.

Ready for publication

Multimodal Medical AI for Chest Disease Classification

ResNet-50 + BioClinicalBERT · IU Chest X-Ray

End-to-end system across 16 pathology categories: 95.20% test accuracy, 0.9921 weighted AUC-ROC. Grad-CAM explainability, severity prediction, and NIH ChestX-ray14 cross-validation.

View code on GitHub →

Journal · Under review

CoNKAN — Kidney Stone Detection in CT Scans

Co-authored · ConvNeXt-Small + KAN attention

ConvNeXt backbone with Kolmogorov-Arnold Network channel attention for CT stone detection. 99.13% accuracy, zero false positives on 346 test images; outperforms ResNet-50, DenseNet-121, VGG-16, and EfficientNet-B2. MixUp, two-stage fine-tuning, ablation study on GitHub.

Broader research focus

  • Multimodal & explainable medical AI — vision + clinical text with Grad-CAM-style interpretability for assistive diagnosis.
  • Efficient video & detection at scale — transformers and YOLO families with reproducible training and validation.
  • Operational CV — traffic analytics, OCR, and annotation-to-deployment pipelines from NCAI lab work.

More experiments and notebooks on GitHub. Preparing MS-oriented applications and open to academic collaborations in computer vision and trustworthy AI.

Skills

Technical skills

Stack I use daily for CV/ML delivery and research prototypes.

CV

Computer vision & deep learning

YOLOv5 / v8 / v11 OpenCV PyTorch TensorFlow MediaPipe MMAction2 VLMs
ML

Data & deployment

Roboflow CVAT Jupyter Kaggle RunPod Streamlit Architectural CV

Languages

Python C++ JavaScript HTML / CSS
+

Medical & OCR

Segmentation Multimodal (ViT / BERT) PaddleOCR Ultrasound / MRI-style CV

Tracking & analytics

Object tracking Motion analysis Traffic analytics

Other

Git / GitHub Urdu · Pashto · English

Projects

Projects

Client deliverables and research-grade builds below. Open-source work and notebooks are listed under “More repositories” and on GitHub.

Client & deliverable projects

Client · Remote sensing & GIS

Flood Hazard & Urban Sprawl Prediction

Islamabad Capital Territory · ML · Python · 2030–2050

End-to-end ML pipeline predicting urbanization and flood hazard from LULC satellite data. Six regression models per land-cover class (LOOCV); vectorized polynomial regression across 35M HEC-RAS pixels (Depth, Velocity, WSE). Near-perfect correlation (r = 0.99) between sprawl and worsening flood risk. Delivered GeoTIFF maps, ArcGIS overlays, notebooks, and technical documentation.

University of Stirling, UK · MSc AI / Big Data

Speech Emotion Recognition

Audio ML · PyTorch · librosa

End-to-end SER system: merged RAVDESS, CREMA-D, TESS, and SAVEE (62,000+ samples), 234 acoustic features per clip (MFCC, Mel, Chroma, Spectral Contrast, Tonnetz). Compared SVM, 1D-CNN, CNN-LSTM, and CNN-LSTM with channel + spatial attention — novel attention variant targeting 93–97% across eight emotions.

Featured open-source projects

AI Traffic Control System

Real-time vehicle detection with signal timing from lane density using YOLOv11.

Pakistani License Plate Detection

YOLOv8 plus PaddleOCR for plate text and CSV export.

Brain Tumor Segmentation

YOLOv11 on MRI-style scans for tumor and dead-cell regions.


Appliance Status Detection

Classify whether appliances are on or off from camera frames with YOLO.

Ambulance & Rickshaw Detection

Multi-class YOLOv11 detectors for emergency and rickshaw traffic.

More repositories

Each line is a one-line summary. To read code, notebooks, or train configs, open the repo on GitHub — everything lives under github.com/daud-shah.

Want the full picture? Visit all repositories on my GitHub profile.

More projects on Github

I love to solve AI problems


GitHub

Contact

Let's Connect

Have a project in mind or want to collaborate on AI research? Let's connect and build something incredible.

Location

Islamabad, Pakistan