About the role
This is not a research post and we will not pretend otherwise. The models we need are forecasting models: how much headroom a customer's cluster has before the next quarter-end batch, which metric series has genuinely gone abnormal versus which one always looks like that on a Monday, and how to compress six months of noisy telemetry into an alert an SRE will actually trust.
The advantage of doing applied ML here is the feedback loop. A forecast that is wrong shows up in a capacity review two weeks later, and an anomaly detector that cries wolf gets switched off by the on-call engineer who is sitting fifteen metres away. You will know quickly whether your work is any good.
You will join the Applied ML pod inside Data & Insights — currently two engineers and a data platform engineer — and work day to day with the reliability team who consume everything you build.
What you'll do
- Build and ship forecasting models for capacity and cost across customer environments, with clear uncertainty bounds.
- Develop anomaly detection over high-cardinality metrics that reduces alert noise rather than adding to it.
- Own your models in production: deployment, monitoring, drift detection and retraining schedules.
- Build the feature and evaluation pipelines so results are reproducible and comparable between quarters.
- Work directly with SREs to define what a useful signal looks like before you optimise for it.
- Write up methods and limitations honestly, including the experiments that did not work.
What we're looking for
- Five or more years in machine learning engineering with models you have taken to production and kept there.
- Strong Python, plus experience with a modern ML stack (PyTorch, scikit-learn, or equivalent) and MLOps tooling.
- Time-series experience — forecasting, seasonality, and the failure modes of both.
- Comfortable with cloud infrastructure and containerised deployment; your models run on Kubernetes like everything else here.
- Business-level English.
Nice to have
- AIOps or observability-domain experience.
- Feature store or streaming feature pipelines.
- Publications or open-source contributions in applied ML.
What you get
- 13th month salary and discretionary bonus
- GPU budget for experiments plus fully funded cloud certifications
- Medical and dental for you and your dependants
- 20 days annual leave, rising to 25
- Hybrid working, two days a week at Science Park
Skills & keywords
Hiu-tung Ma
Applied ML Lead · reviews applications personally
Listing ID JOB-TECH-009 · Closes 16 Oct 2026 · HKjobs never asks candidates to pay a fee. Report this listing