Skip to content
Hamidreza Saghir by Hamidreza Saghir

About

Principal Applied Scientist working on agentic AI, LLM evaluation, and production ML systems.

I’m Hamidreza Saghir, a Principal Applied Scientist at Microsoft working on agentic AI, LLM evaluation, and production ML systems.

My recent work focuses on the machinery around agents: planning loops, tool use, context and memory, evaluation, provenance, replay, and CI-gated quality systems. At Microsoft, I established evaluation infrastructure adopted across 20+ Security Copilot skill teams, built systematic evaluation for the Defender investigation agent, and improved agent orchestration performance by 25%.

Before Microsoft, I led multilingual entity intelligence work at X, owned the scientific roadmap for Alexa Automotive edge NLU at Amazon, and worked on NLP and fraud detection at Borealis AI. I have a Ph.D. from the University of Toronto and publications at ACL, COLM, InterSpeech, WWW, IEEE TASLP, and other venues.

This site is where I write about machine learning, agent systems, evaluation, security, and the places where different technical ideas turn out to share the same structure.

Selected work

  • Looplet: creator of an open-source Python harness for observable, testable tool-calling agents with capture, replay, policy hooks, and outcome-grounded evaluation.
  • Security Copilot: built shared evaluation infrastructure adopted across 20+ skill teams and took it from design through public-preview launch.
  • Microsoft Defender: built the first systematic LLM-as-judge evaluation pipeline for the investigation agent.
  • X / Twitter: shipped multilingual entity intelligence and entity-aware topic classification to all users.
  • Amazon Alexa: owned the scientific roadmap for Automotive edge NLU and shipped an 80% model-footprint reduction on OEM devices.
  • Borealis AI / RBC: developed NLP models and semi-supervised fraud detection from unlabeled transaction data.

Roles I am best matched for

  • Senior Staff, Principal, or Senior Principal roles in agentic AI, LLM systems, evaluation infrastructure, ML systems, NLP, retrieval, or developer tools.
  • Technical leadership roles where architecture, measurement, and production delivery matter more than demos.
  • Selective Head of AI or Director roles only when the mandate includes real technical authority, strong execution ownership, and meaningful compensation.

Elsewhere

The best way to reach me is email. I try to reply, though I may be slow.