AWS OPERATIONS
What should your AWS managed services agreement cover?
A practical guide to defining operating scope, agreeing responsibilities, and checking that day-to-day AWS work has an owner.
1. Define the systems and outcomes in scope
For a Thai SME expanding across ASEAN, an AWS operations agreement should begin with the business systems that need care: an ordering portal, a warehouse application, or a reporting platform. List the accounts, environments, dependencies, and users involved. A phrase such as “manage our cloud” leaves too much room for different expectations.
In this guide, managed services means an agreed service scope for operating your AWS environment, such as a scope discussed with BlissJunction. Begin with business impact: what stops when a system is unavailable, who is affected, and which trading periods require special preparation? Use those answers to set priorities.
- Name an internal business owner for each application and identify its technical contact.
- List included environments and explicitly record third-party systems and excluded work.
- Agree how new countries, applications, or AWS accounts enter the service scope.
2. Make responsibilities visible
AWS describes security as a shared responsibility. AWS protects the underlying cloud infrastructure; customers retain responsibilities that depend on the services and configuration they use. Hiring an operations provider adds another working relationship, so write down which customer tasks the provider will perform.
Build a simple matrix with one row per activity and columns for the person doing the work, the approver, and the people informed. Include application code and data decisions as well as infrastructure. Walk through a realistic issue together to expose gaps between teams before an incident does.
- For access changes, identify who requests, approves, implements, and reviews access.
- For patching, separate infrastructure work from application compatibility testing and release approval.
- For AWS support cases, define who opens the case and supplies application evidence.
3. Specify the recurring work and its evidence
Useful scope describes both an activity and how completion is demonstrated. “Monitoring” should identify the symptoms that matter to users, the alerts being watched, and the action expected when an alert fires. A long list of metrics is less useful if nobody owns the response.
Ask for a recurring work schedule covering agreed maintenance, access reviews, capacity, and cost review. Document routine procedures in runbooks, including prerequisites, permissions, error handling, and escalation. For backups, agree retention and test recovery against business objectives: acceptable downtime and acceptable data loss.
- Record backup results and periodically test that restored data is usable.
- Keep maintenance records, unresolved exceptions, and the next planned action together.
- Review spending changes with an owner who can approve configuration or usage changes.
4. Agree how incidents and changes are handled
An alert, an incident, and a change request need different handling. Define incident severity by business impact, then agree contact channels, service hours, escalation paths, and update expectations. Record response and recovery expectations separately; acknowledging a problem does not mean the service has recovered.
For planned work, specify who assesses risk, approves the change, tests the result, and decides whether to roll back. Include a route for urgent changes when normal approval is unavailable. Businesses serving several ASEAN markets should state the time zone for maintenance windows and the language used for operational updates.
- Test the contact tree and keep an alternate approver available.
- Identify actions the provider may take immediately and actions requiring approval.
- Review significant incidents for causes, follow-up owners, and due dates.
5. Keep the service reviewable and transferable
A service review should help a business owner make decisions. Review recurring incidents, overdue work, restore-test findings, changes in spending, and risks that need business acceptance. Pair each open issue with an owner and a date. Add context when a quieter month reflects reduced usage rather than an operational improvement.
Agree handover requirements at the start. Your business needs continued access to its AWS accounts, operational records, diagrams, and agreed configuration documentation. Establish how privileged access is removed and how unfinished work is transferred when personnel or suppliers change.
- Choose a review cadence that matches system importance and the pace of change.
- Ask for evidence of completed work, alongside recommendations requiring a decision.
- Update the responsibility matrix whenever the application or team changes.
6. Illustrative scenario: an ASEAN ordering portal
Consider a fictional Thai distributor launching a dealer portal for Malaysia and Vietnam. Its developers own application releases; an operations provider handles an agreed AWS infrastructure scope; the sales operations manager owns the order process. This is a planning example, not a customer case study or a promised service package.
Before launch, the team rehearses a failed order submission. Infrastructure checks may be healthy while an application validation rule rejects orders. The exercise establishes who gathers logs, who investigates code, who updates sales teams, and who authorizes rollback. The agreement is ready when these handoffs can be followed without guessing.
- Bring a system inventory, business calendar, and named decision-makers to the scoping discussion.
- Use one routine task, one incident, and one recovery exercise to test the proposed scope.
- Confirm the remaining gaps and assign them before the service begins.
