Mirendil · Member of Technical Staff, Model Evaluation

Know whether the model is actually getting better in ways that matter.

Mohamed A M Elansary, PhD — scientific multimodel evaluation, uncertainty quantification, and production agent evaluation sets for measurement beyond a single benchmark score.

Model evaluationUncertainty quantificationAgent evalsScientific ML

Evaluation under uncertainty

  • Six-plus years of multimodel, multi-basin forecast experiments across hydroclimates on Linux/HPC.
  • Compared statistical and physically based stacks, quantified uncertainty, and reported regime-dependent failure modes rather than a single flattering score.
  • That is the measurement analogue of asking whether a frontier system is actually improving along realistic axes, beyond standard benchmarks.

Agent evaluation sets

  • Production GPT, Claude, and Gemini agent workflows with retrieval, routing, tenant isolation, provenance, and regression evaluation sets at Vertexium, including a multi-tenant conversational receptionist.
  • That maps to automated eval / regression-detection practice and inspecting multi-step behavior. It is not Mirendil-internal eval-framework ownership.
  • Scientific/HPC execution, data pipelines, and clear technical writing support owning frameworks, pipelines, and tooling that surface signal quickly.

Proposed first contribution

For one capability axis already in flight, define what “better in ways we care about” means as observable evidence versus a score that is easy to move. Write a small failure taxonomy: metric movement without a causal behavioral change; a regression suite that passes on a demo slice and fails under regime shift; observability that looks informative but does not predict held-out failure; uncertainty that collapses when observations are imperfect. Stand up a small evaluation set with provenance, compare simple baselines, attach uncertainty, and write a clear report that can hand signal to post-training and RL partners before expanding the framework. This is a proposed measurement approach, not a claim of prior Mirendil-internal work, invented metrics, safety research, or RLHF.

Honest fit boundary

Mirendil-specific product/eval stack is a stretch. I have not owned Mirendil-internal evaluation frameworks, instrumented Mirendil training runs, partnered with Mirendil post-training/RL teams, published frontier LLM benchmark suites, or claimed Mirendil-internal work, and I do not invent metrics, safety research, or RLHF. The credible contribution is scientific multimodel evaluation, uncertainty quantification, production agent evaluation harnesses, scientific/HPC rigor, data pipelines, and clear technical writing.

Role and location

Member of Technical Staff, Model Evaluation · San Francisco · OnSite. Willing to relocate to San Francisco with a relocation package. Remote work is not asserted.

Posting compensation: “We offer a base salary of $300,000–$400,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.” · Official role posting