WarburgAI migrates a GPU-heavy FinServ AI workload to Amazon SageMaker
Shared anonymously — the customer’s name is held by VeUP and available on request.
GPU is the most expensive line on an AI company’s cloud bill — and the easiest to lose control of. VeUP moved this firm’s AI workload off another cloud and onto Amazon SageMaker, reserved its GPU through Capacity Blocks for ML, shifted inference to AWS Graviton, and wrapped the whole estate in a multi-account FinOps structure that made the spend visible, predictable, and governed.
The challenge
The firm — a Dubai-based, AI-driven financial-services company — ran its GPU/ML workload on a different cloud and wanted to scale on the AWS AI/ML ecosystem. But the move only made sense if the spend came with it under discipline. On-demand GPU is the most expensive way to run a bursty ML workload; a single-account billing posture meant nobody below root could see what anything cost; and a Marketplace-vs-direct spend mis-attribution had quietly distorted the whole cost picture.
The solution
VeUP ran the migration through all three phases — assess, mobilize, migrate — and landed the AI workload on Amazon SageMaker. Two cost decisions carry the design: GPU capacity for training and inference is reserved through Capacity Blocks for ML, trading unpredictable on-demand pricing for a known number, and inference runs on AWS Graviton for better price-performance. Underneath, a multi-account payer/child structure with Cost Explorer and a CUR-ingest cost platform gives the firm scoped, non-root cost visibility and a monthly optimization cadence — with VeUP running the engagement as managed billing and FinOps.
Production outcomes
| KPI | Result |
|---|---|
| Production outcomes | The AI workload runs in production on AWS. FinOps corrected the Marketplace-vs-direct spend mis-attribution that had distorted the cost picture, then shifted the compute mix toward Capacity Blocks for ML and Graviton inference — reducing the effective run cost of the workload, tracked continuously in CUR dashboards. |
| Timeline | Migration kicked off in mid-December 2024; the workload was live on AWS by the end of January 2025 — about six weeks, cutover included. VeUP has run managed billing and FinOps for the firm since. |
| Cost posture | Cost engineering is the heart of the engagement: a monthly FinOps cadence works the levers — Capacity Blocks for ML over on-demand GPU, Graviton for inference — with a CUR-ingest cost platform and Cost Explorer providing visibility across the multi-account payer/child structure. |
| Lessons & continuation | For a GPU-heavy FinServ AI workload, the controlling cost levers are reserving GPU via Capacity Blocks for ML (not on-demand) and shifting inference to Graviton; a multi-account payer/child + CUR structure is the precondition for scoped, non-root cost discipline; correcting a Marketplace-vs-direct spend mis-attribution before optimizing avoids modeling against a distorted baseline. |
Architecture

Where it started
Assessed baseline · pre-migration source cloudFinancial-services AI · Middle East · workload on another cloud providerGPU/ML training and a live production model-serving loop running on the source cloud, with a Redis-style cache and object-backed data tier feeding the pipeline.
No reserved capacity — the most expensive pattern for a bursty ML workload, with uncontrolled on-demand spend driving run cost.
Single-account billing with Marketplace-vs-direct mis-attribution — no scoped view of what the workload truly cost.
A single serving loop with no multi-AZ posture for the live production model-serving path.
No access to Capacity Blocks for ML or Graviton economics while the workload stayed on the source cloud.
The source-cloud estate as assessed before the migration — the serving loop was later re-platformed onto Amazon SageMaker.