Someone has to own
the pager
Go-live is the start of the expensive part. Our managed service covers the run — monitoring, patching, incident response and the continuous improvement backlog — without becoming a dependency you cannot leave.
A managed service that keeps making the platform better
Most managed services are priced to keep things exactly as they are, because change creates risk for the provider. Ours includes an improvement backlog, because a platform that does not improve is depreciating.
Run
Monitoring against agreed SLOs, patching, backup verification, certificate and dependency currency, and incident response with a named escalation path.
Report
Monthly SLO attainment, security posture, cost trend and incident review — written for your leadership, not exported from a tool.
Improve
A standing backlog of reliability, security and cost work, prioritised with you each quarter and delivered within the retainer.
Hand back
Everything in your accounts, in version control, with runbooks. Exit assistance is written into the agreement, not negotiated at the end.
Three levels of cover
Most clients start at Core and move to Critical when a workload becomes genuinely business-critical rather than important.
| Essential | Core | Critical | |
|---|---|---|---|
| Coverage | Business hours AEST | Business hours L2, 24×7 L3 | 24×7 L2 and L3 |
| P1 response | 2 hours | 30 minutes | 15 minutes |
| Monitoring | Infrastructure | Infrastructure and application SLOs | Full stack with synthetic checks |
| Patching | Monthly cycle | Monthly, with emergency out-of-band | Continuous, immutable rebuild |
| Reporting | Monthly summary | Monthly SLO, security and cost | Monthly plus quarterly architecture review |
| Improvement backlog | — | Included, quarterly prioritised | Included, with dedicated capacity |
| DR exercises | Annual | Semi-annual | Quarterly, with evidence pack |
Response targets are measured from alert acknowledgement. Full service-level definitions form part of the service schedule.
Managing the AI systems too
Production AI needs operating in ways traditional monitoring does not cover. Our Critical tier extends to the systems built on top of the platform.
Quality monitoring
Evaluation suites run on a schedule, not only at deployment, so retrieval decay surfaces before users lose trust.
Cost per interaction
Inference spend tracked per use case against agreed ceilings, with alerting on drift.
Drift and freshness
Index freshness, embedding coverage and model performance tracked as first-class service metrics.
Guardrail review
Blocked and flagged interactions reviewed monthly, feeding both the evaluation set and the guardrail configuration.
Taking over a platform we did not build
Roughly half our managed clients came to us with an existing environment. The onboarding is deliberate and it takes six weeks.
Discovery and risk assessment
Full inventory, dependency mapping, backup verification and a candid risk register. We tell you what we found, including the parts that are uncomfortable.
Stabilisation
Monitoring and alerting to our standard, access and secrets brought into order, and any critical gaps closed before we accept the pager.
Shadow operation
We run alongside your team, taking alerts and handling incidents with your people watching, so the handover is observed rather than assumed.
Service commencement
SLOs agreed, runbooks published, escalation paths tested. The improvement backlog starts with whatever the risk register surfaced.
Tell us what wakes your team up
The alerts, the manual steps, the person who cannot take leave. Those are the things a managed service should remove, and the first conversation is about exactly that.