Back to
Metal
20 Moves that take a company off AWS, Google Cloud or Azure and onto dedicated servers you rent or own in colocation.
Written for a company whose platform is two or three engineers' work and whose cloud bill is around $100,000 a month. Keep the team, bring your platform skills, and let remote hands handle the physical work. Rent to preserve cash; colocate to own the capacity.
The whole plan, on one page
5 stages, run in order, and every dependency points at a lower number. 18 of the 20 Moves plan for zero downtime; the reference cutovers total 25 minutes, and 1 of them cannot be undone. These estimates depend on each Move's prerequisites and a timed rehearsal.
One bar a Move, as tall as the days it takes, in the colour of its stage.
- Stage 1DecideWork out whether to do it at all Done when You know what you spend, what you would spend instead, and whether it is worth it. - 01The bill, and the three lines that are most of itThe invoice records what you were charged, not what you use. Read it to the line, find the three that are most of it, then start the clock on working set.Effort 3 days Risk Low Cutover no downtime
- 02What you actually runThe estate feels infinite because nobody has written it down. Give every workload a row and an owner, then keep its resources underneath it.Effort 4 days Risk Low Cutover no downtime
- 03The number that decides itThree columns, the same operations team and five years of costs. Put colocation and rented metal beside the cloud bill and choose the margin worth moving for.Effort 2 days Risk Medium Cutover no downtime
- 04The three things you keep rentingThree capabilities do not come home: the edge, outbound mail and volumetric scrubbing. Naming them now keeps the plan honest.Effort 2 days Risk Low Cutover no downtime
- Stage 2BuyOrder the hardware and sign the space Done when The machines are on order and the cage is signed. - 05From rented vCPUs to cores you ownA cloud vCPU is a scheduling unit. Measure the workload before buying cores, then leave room for the guests and a failed host.Effort 3 days Risk Medium Cutover no downtime
- 06Sixteen machines, and the two on the shelfSixteen active nodes and two reserved replacements, bought for the shelf or included in the rental quote, because iron does not autoscale.Effort 3 days Risk High Cutover no downtime
- 07A cage, not a data centreThe facility supplies power, cooling, connectivity and remote hands. Your team operates the platform; rented metal includes the room and the machines.Effort 4 days Risk High Cutover no downtime
- 08The order, and the weeks you cannot compressEverything with a lead time goes on one order this afternoon, so the queues run beside each other instead of one after another.Effort 3 days Risk Medium Cutover no downtime
- Stage 3BuildTurn the boxes into a cluster Done when A cluster that could take production traffic, and never has. - 09Racking dayProve remote console access to the whole batch, then shelve the spares. Rails, power and labels are the work; recovery from outside the building is the proof.Effort 6 days Risk Medium Cutover no downtime
- 10The network, and the way back in when it breaksAddress ranges written down before anything is configured, two switches either of which may fail, and a way in that depends on neither.Effort 4 days Risk High Cutover no downtime
- 11The platform your workloads needCloud VMs become KVM guests on Proxmox VE. Containers keep Kubernetes. Owning the hardware does not require changing the application's packaging.Effort 4 days Risk High Cutover no downtime
- 12Disks: what goes local, what goes on CephChoose who owns each physical disk. Proxmox supplies VM datastores; Rook supplies storage on the direct Talos path. Replication belongs in one layer.Effort 6 days Risk High Cutover no downtime
- Stage 4MoveMove the app, then the data Done when Everything runs on your machines, and the cloud copy is still warm. - 13Images, secrets and one-command deploysRegistry, secrets and deployments are proved together in staging, before customers depend on the new platform.Effort 7 days Risk Medium Cutover no downtime
- 14The first service, end to endThe first service runs in a VM or a container on your hardware, with its data still in the cloud and a traffic weight as the way back.Effort 20 days Risk Medium Cutover no downtime
- 15Buckets, cache and queuesCopying bytes is the easy part. A live move also needs an ordered record of changes, retryable publishing and consumers that recognise the same job twice.Effort 8 days Risk High Cutover no downtime
- 16Postgres, the one that mattersFifteen minutes of stopped writes, at the end of a week in which the new cluster was caught up, backed up, and restored once on purpose.Effort 12 days Risk High Cutover 15 min down
- Stage 5RunCut the traffic over, and keep it alive Done when Users reach your machines, you can carry it at 03:00, and the cloud bill is zero. - 17The front doorThe new edge goes up beside the one carrying users, with its own address and certificates, proved from a host file entry while real traffic goes elsewhere.Effort 4 days Risk Medium Cutover no downtime
- 18Go-live, and how you abortShift traffic in measured steps, with both edges reaching the same data. Lower DNS lifetimes ahead of time and keep the old edge available for a week.Effort 3 days Risk High Cutover 10 min down
- 19Backups you have restored, and the pagerVM backups, cluster manifests and a continuous database archive, copied beyond the rack — then the drill that measures recovery time.Effort 5 days Risk High Cutover no downtime
- 20Closing the accountThe estate moved a month ago. The account did not: a dormant organisation still bills, still holds credentials, and still lets someone start an instance in it.Effort 4 days Risk High Cutover no downtime
Start where it hurts
Nobody sits down at nine in the morning thinking in stages. Find the line that sounds like your week and go straight to the Move.
- The bill went up and nobody can say why01 02 03
- Somebody has asked what would happen if we just left03 04 07
- Egress is now one of the largest lines on the invoice01 14 18
- The database costs more than the engineers who query it05 12 16
- We have no idea what we are actually running02 13 20
- We are being asked to do this and we have two engineers03 06 19
- The hardware sounds like a full-time job nobody has07 09 11
- Something has to move this quarter and nothing may break14 17 18
- We want remote hands to handle the hardware calls07 19 06
- We were told we cannot leave, and nobody has checked04 15 20
Why leave at all
Colocation and rented dedicated servers put more of your infrastructure budget into capacity you can use. Rent the machines to preserve cash, or own them in a colocation facility for lower running costs and control over the hardware. In this reference model, colocation saves $4,886,320 over five years, with the same operations hours as the cloud. Your team brings its skills; the provider supplies the building and physical support.
The evidence to collect before claiming a win
- UptimeMeasure from outside both estates, with the same service boundary and observation window. Count site failures and planned downtime.
- LatencyReplay representative traffic on both platforms. Compare throughput and tail latency at the same offered load.
- RepairsRecord physical interventions, time to recovery and remote-hands charges. A response-time promise is not a repair-time measurement.
- HoursRecord recurring platform work on both sides. Separate migration labour from steady-state operations and on-call coverage.
The financial figures are a worked model. They do not establish an availability, performance or staffing result for your estate.
- 01The bill stops being a percentage of your growthThis is the one that matters and the rest are consequences of it. A cloud bill is a toll on activity: more users, more bytes, more bill, for ever, at a margin somebody else sets and can change. A rack or a dedicated-server contract buys capacity at a predictable price. While your workload fits that capacity and your power and traffic allowances, more users need not mean more infrastructure rent. Add machines when demand calls for them. Over five years the difference here is $4,886,320, and the machines are still working at the end of it. Check itTake last quarter’s invoices and plot the total against your own usage metric. If the line is flat you have nothing to gain here. If it tracks your growth, that slope is what you are buying out of — Move 01.
- 02Egress stops being a tax on your own trafficMoving your own data to your own users is where the margin lives, and it is the line to price explicitly: $7,980 a month for 100 TB of AWS internet egress at the model's US East list rates. Facility transit replaces that meter with a bandwidth commitment and possible overage charges. Compare the actual traffic pattern and contract on both sides. Check itFind the data-transfer line on last month’s bill. Divide it by your egress in terabytes. Compare that effective rate with a transit quote and the edge services you retain — Move 01, then Move 18.
- 03The machines are yours, so the performance is yoursA vCPU can be an SMT thread or a whole physical core; the instance family decides. Dedicated hardware gives you control over placement and contention, and local NVMe changes the storage path. Neither guarantees faster queries: CPU generation, memory, disks and workload still decide the result. Check itCompare your ninety-ninth percentile against your median for a week. If the gap is wide, investigate contention and application behaviour. Replay the same load on real hardware in Move 14 before attributing the gap to the cloud.
- 04You control the upgrade windowA managed service is somebody else’s roadmap running inside your product. Instance families are retired, versions go end-of-life on a date you did not pick, a control plane upgrades on its own release channel, and the longest you can defer an upgrade depends on the service and support policy. On your own platform you schedule the upgrade, but software and firmware still reach end of support. Budget patching and replacement; ownership does not make an unsupported version safe. Check itCount the forced upgrades and deprecation notices you have absorbed in the last two years, and what each cost in engineer-days. That is a recurring bill nobody invoices you for.
- 05Your team runs the platform; remote hands look after the rackYour cloud-ops engineers already deploy, monitor, patch and recover services. Those skills move with the workload. Colocation remote hands carry out disk swaps, cabling and power cycles under your runbooks; a rented-metal provider maintains its hardware under the support agreement. Physical support is already priced in facility charges or rent. The model retains the existing team, hours and salary, without counting potential savings from lower-cost on-premises roles. Check itAgree the physical tasks, coverage and response times with the provider, then rehearse an incident together. Measure your team’s hours before and after the move — Moves 07 and 19.
Choose how you move
Both routes can lower the bill. Choose the purchasing and support arrangements that fit your cash flow and the team you already have.
- 01Rent first, or own in colocationRented metal lets you move without buying the servers upfront. Colocation lets you own the hardware and spread its purchase cost over years of use. Compare both cash flows in Move 03 and choose the route in Move 07.
- 02A timetable that fits the routeAvailable rented servers can avoid the hardware purchase and facility setup lead times. The colocation plan includes ordering machines, signing space and waiting for circuits: 107 person-days of work across 41 weeks. Both routes keep the migration rehearsals and rollback windows.
- 03The existing team, with no added salary in the modelCloud ops transitions into on-prem ops. Remote hands handles the physical interventions; your engineers keep responsibility for the platform and application. The model budgets 160 hours a month in every option, with remote hands in the facility bill. On-premises roles can cost less than cloud specialist roles; that potential saving is excluded by keeping the same salary. Validate local rates and hours. Migration work is counted separately.
- 04Three things that never come homeA content delivery network, outbound mail deliverability and scrubbing at the edge are businesses other people run better than you will. Move 04 tells you to keep paying for all three. These and the services retained in later Moves total $9,180 a month in the reference model.
When to stay exactly where you are
- Your load is spiky or seasonal and genuinely scales to near nothing between the spikes. Elasticity is the one thing a rack cannot do, and paying for the peak all month is how owning becomes the expensive answer.
- You lean hard on managed services that have no real equivalent - a serverless database, a hosted stream, a workflow engine - and replacing them is a rewrite rather than a migration. Move 02 will tell you this in an afternoon.
- You have no team to own the platform and no managed operations partner. Remote hands covers physical work; application recovery and platform decisions still need an owner. An existing cloud-ops team can transition into that role.
- Your effective cloud bill is too small to cover migration and the capacity you need. Price rented metal as well as a colocation rack: renting can suit a smaller estate. Move 03 compares your own quotes and hours, so the decision rests on your workload rather than a spending threshold borrowed from another company.
If those reasons to stay do not apply and the arithmetic in Move 03 clears, the remaining Moves provide a staged plan, with a stated rollback limit for each job.
How long it takes
The labour is about 107 person-days. With 2 engineers it lands in roughly 41 weeks, and the shortest it could possibly take, with as many people as you care to put on it, is 40 weeks under these estimates. Limited staffing adds contention for engineers. Even unlimited staffing cannot remove the ordered work, circuit orders, hardware lead times and observation windows on the critical path. Calendar waits consume no engineer-days.
One bar per Move, drawn from each Move’s own effort figure and the dependency graph rather than from a plan somebody typed. The outlined bars are the critical path.