Ramp , a finance automation platform processing billions of dollars in spend decisions, runs GPU-powered AI inference continuously on Amazon Elastic Container Service (Amazon ECS) . As the company scaled its machine learning workloads, managing the underlying Amazon Elastic Compute Cloud (Amazon EC2) fleet, Auto Scaling groups, launch templates, custom Amazon Machine Image (AMI) patching, and monitoring scripts demanded disproportionate engineering effort.

Amazon ECS Managed Instances significantly reduced that overhead, bringing Ramp’s GPU infrastructure into full operational parity with the rest of their ECS infrastructure. This post walks through how Ramp’s infrastructure team made that transition: the architecture pattern, implementation details, and what they learned along the way. Ramp is a finance automation platform trusted by more than 70,000 businesses (source: ramp. The platform helps them save time and money. The platform combines corporate cards, expense management, bill pay, procurement, travel, treasury, and accounting automation in a single system.

Since its founding, Ramp’s mission has been to help businesses get more out of every dollar and every hour. Advanced analytics and machine learning (ML) have powered that mission from the start. Ramp has used ML for years to provide recommendations, extract insights, and combat fraud. Today, Ramp is evolving from basic automation to an intelligent finance operations platform that manages the entire spend lifecycle, from purchase order requests through month-end close. In October 2025 alone, Ramp’s ML models made more than 26 million decisions across over $10 billion in spend, according to internal Ramp metrics.

It prevented hundreds of millions of dollars in out-of-policy transactions and flagged fraudulent invoices in real time. Delivering AI at that scale, reliably and in real time, requires substantial investment in GPU-accelerated machine learning infrastructure. Behind automated transaction categorizations, merchant matches, semantic search results, and receipt extractions, GPU inference workloads run continuously in production.