About
Principal Applied Scientist working on agentic AI, LLM evaluation, and production ML systems.
I’m Hamidreza Saghir, a Principal Applied Scientist at Microsoft working on agentic AI, LLM evaluation, and production ML systems.
My recent work focuses on the machinery around agents: planning loops, tool use, context and memory, evaluation, provenance, replay, and CI-gated quality systems. At Microsoft, I established evaluation infrastructure adopted across 20+ Security Copilot skill teams, built systematic evaluation for the Defender investigation agent, and improved agent orchestration performance by 25%.
Before Microsoft, I led multilingual entity intelligence work at X, owned the scientific roadmap for Alexa Automotive edge NLU at Amazon, and worked on NLP and fraud detection at Borealis AI. I have a Ph.D. from the University of Toronto and publications at ACL, COLM, InterSpeech, WWW, IEEE TASLP, and other venues.
This site is where I write about machine learning, agent systems, evaluation, security, and the places where different technical ideas turn out to share the same structure.
Selected work
- Looplet: creator of an open-source Python harness for observable, testable tool-calling agents with capture, replay, policy hooks, and outcome-grounded evaluation.
- Security Copilot: built shared evaluation infrastructure adopted across 20+ skill teams and took it from design through public-preview launch.
- Microsoft Defender: built the first systematic LLM-as-judge evaluation pipeline for the investigation agent.
- X / Twitter: shipped multilingual entity intelligence and entity-aware topic classification to all users.
- Amazon Alexa: owned the scientific roadmap for Automotive edge NLU and shipped an 80% model-footprint reduction on OEM devices.
- Borealis AI / RBC: developed NLP models and semi-supervised fraud detection from unlabeled transaction data.
Roles I am best matched for
- Senior Staff, Principal, or Senior Principal roles in agentic AI, LLM systems, evaluation infrastructure, ML systems, NLP, retrieval, or developer tools.
- Technical leadership roles where architecture, measurement, and production delivery matter more than demos.
- Selective Head of AI or Director roles only when the mandate includes real technical authority, strong execution ownership, and meaningful compensation.
Elsewhere
- ML interview prep: mlmentorship.com - senior ML/AI interview questions, guides, and concept notes.
- Resume: PDF
- Looplet: GitHub and docs
- Email: saghir.hr@gmail.com
- GitHub: @hsaghir
- LinkedIn: hamidrezasaghir
- Google Scholar: publications
The best way to reach me is email. I try to reply, though I may be slow.