Cloud Platforms for AI: AWS, Azure & GCP
After this lesson, you will be able to:
- Compare the big three cloud AI platforms (AWS SageMaker, Azure ML, GCP Vertex AI) for training, serving, and managing models
- Pick the right cloud provider for your specific workload based on cost, features, and what your team already uses
- Design a cloud setup for fine-tuning a foundation model using managed services from any of the three providers
Before You Start
#The Three Giants
Cloud platforms can feel overwhelming with their hundreds of services, but do not worry — by the end of this lesson, you will know exactly which services matter for ML and how to choose between providers. Most of the decision comes down to practical factors, not feature checklists.
You need to fine-tune a 7B parameter model on 100K training examples. Which cloud service would you choose?
The answer is D. All three clouds can handle this workload. The deciding factors are rarely technical — they are organizational. Do you have an existing AWS account with credits? Is your company on Microsoft Enterprise Agreement? Does your team have GCP experience? Start from your constraints, not from a feature comparison chart.
#AWS for AI: The Market Leader
AWS holds ~32% of the cloud market and has the broadest AI service portfolio. Its strength is sheer breadth: if an AI service exists, AWS probably offers it.
#Amazon SageMaker
SageMaker is AWS's end-to-end ML platform. Think of it as a managed Jupyter notebook that grew into an entire ML factory.
Core capabilities
| Feature | What It Does | When to Use |
|---|---|---|
| SageMaker Studio | Managed IDE (JupyterLab) with built-in experiment tracking | Daily ML development |
| Training Jobs | Managed distributed training on p4d/p5 (A100/H100) instances | Fine-tuning, pre-training |
| Endpoints | Real-time inference with auto-scaling | Production model serving |
| Batch Transform | Batch inference on large datasets | Offline scoring, nightly jobs |
| Pipelines | ML workflow orchestration (DAGs) | Automated retraining |
| JumpStart | Pre-trained model hub (Llama, Stable Diffusion, etc.) | Quick deployment of foundation models |
| Feature Store | Centralized feature management | Shared feature engineering |
Try it! Go to the free tiers of all three cloud providers (AWS Free Tier, Azure Free Account, GCP Free Trial) and navigate to their AI/ML service pages. Compare how each one organizes its ML tools. You will immediately see the personality differences described above — AWS's massive list, Azure's enterprise polish, and GCP's developer-friendly simplicity.
Training a model on SageMaker
# SageMaker training job — the key abstraction is the Estimator
import sagemaker
from sagemaker.huggingface import HuggingFace
estimator = HuggingFace(
entry_point="train.py", # Your training script
instance_type="ml.p4d.24xlarge", # 8x A100 GPUs
instance_count=1,
transformers_version="4.37",
pytorch_version="2.1",
py_version="py310",
role=sagemaker.get_execution_role(),
hyperparameters={
"model_name": "meta-llama/Llama-2-7b-hf",
"epochs": 3,
"batch_size": 4,
"learning_rate": 2e-5,
},
)
# This launches a managed training job — you pay only while it runs
estimator.fit({"train": "s3://my-bucket/training-data/"})#Amazon Bedrock
Bedrock is AWS's managed LLM service — think "API gateway to foundation models." You do not train or host anything. You send prompts, get responses, and pay per token.
- Data stays within your AWS VPC (privacy/compliance)
- Unified billing through AWS
- Fine-tuning and RAG built in (Knowledge Bases)
- Guardrails for content filtering
- Model evaluation tools
#AWS Data Layer for AI
- S3: Object storage for training data, model artifacts, logs. The de facto standard.
- Glue: ETL for preparing training datasets at scale.
- Athena: SQL queries on S3 data without loading it into a database.
- Redshift: Data warehouse for structured analytics.
#Azure for AI: The Enterprise Choice
Azure holds ~23% of the cloud market and has a unique advantage: the deepest integration with OpenAI models. If your company uses Microsoft 365, Teams, or has an Enterprise Agreement, Azure is the path of least resistance.
#Azure Machine Learning
Azure ML is Microsoft's end-to-end ML platform. Its key differentiator is tight integration with the Microsoft ecosystem (VS Code, GitHub, Active Directory).
Core capabilities
| Feature | What It Does | When to Use |
|---|---|---|
| Workspace | Central hub for experiments, models, data, compute | All ML projects |
| Compute Instances | Managed VMs for development (NC, ND series) | Exploration, prototyping |
| Compute Clusters | Auto-scaling GPU clusters for training | Distributed training |
| Managed Endpoints | Real-time and batch inference with auto-scaling | Production serving |
| Pipelines | ML workflow orchestration (similar to SageMaker Pipelines) | Automated retraining |
| Model Registry | Version control and lineage tracking for models | Model governance |
| Prompt Flow | Visual tool for building LLM applications | RAG, prompt engineering |
Training a model on Azure ML
# Azure ML training job using the v2 SDK
from azure.ai.ml import MLClient, command, Input
from azure.identity import DefaultAzureCredential
ml_client = MLClient(
DefaultAzureCredential(),
subscription_id="your-sub-id",
resource_group_name="my-rg",
workspace_name="my-workspace",
)
# Define the training job
training_job = command(
code="./src", # Local training script directory
command="python train.py "
"--model_name ${{inputs.model_name}} "
"--epochs 3 --batch_size 4 --lr 2e-5",
inputs={
"model_name": "meta-llama/Llama-2-7b-hf",
},
environment="AzureML-pytorch-2.1-cuda12@latest",
compute="gpu-cluster", # NC A100 v4 cluster
instance_count=1,
)
# Submit — Azure ML provisions compute, runs training, saves outputs
returned_job = ml_client.jobs.create_or_update(training_job)#Azure OpenAI Service
This is Azure's killer feature for AI. You get OpenAI models (GPT-4, GPT-4o, DALL-E, Whisper) hosted within Azure's infrastructure, with enterprise security.
- Data does not leave your Azure tenant (critical for regulated industries)
- Virtual network integration, private endpoints
- Content filtering and abuse monitoring built in
- SLA guarantees (99.9% uptime)
- Pay with existing Azure commitment
#Azure Data Layer for AI
- Blob Storage: Azure's equivalent of S3 for training data and artifacts.
- Cosmos DB: Globally distributed NoSQL database, excellent for AI application state and vector search.
- Synapse Analytics: Unified analytics platform (data warehouse + Spark + SQL).
- AI Search: Vector and hybrid search for RAG applications.
#GCP for AI: The Research-Driven Choice
GCP holds ~11% of the cloud market but punches above its weight in AI. Google invented the Transformer, built TensorFlow, and designed custom TPU hardware. If cutting-edge AI research matters to you, GCP often gets new capabilities first.
#Vertex AI
Vertex AI is Google's unified ML platform. Its differentiator is the deepest integration with Google's AI research — Gemini models, TPU hardware, and research-grade tools.
Core capabilities
| Feature | What It Does | When to Use |
|---|---|---|
| Workbench | Managed JupyterLab with GPU/TPU support | ML development |
| Training | Custom and AutoML training on GPUs or TPUs | Model training |
| Prediction | Online and batch prediction endpoints | Production serving |
| Pipelines | Kubeflow-based ML workflow orchestration | Automated ML workflows |
| Model Garden | Pre-trained model hub (Gemini, Llama, Stable Diffusion) | Foundation model deployment |
| Feature Store | Centralized feature management | Shared features |
| Generative AI Studio | Prompt design and tuning for Gemini models | LLM application development |
Training a model on Vertex AI
# Vertex AI custom training job
from google.cloud import aiplatform
aiplatform.init(project="my-project", location="us-central1")
# Define a custom training job
job = aiplatform.CustomContainerTrainingJob(
display_name="llama-7b-finetune",
container_uri="us-docker.pkg.dev/vertex-ai/training/pytorch-gpu.2-1:latest",
command=["python", "train.py"],
args=[
"--model_name", "meta-llama/Llama-2-7b-hf",
"--epochs", "3",
"--batch_size", "4",
"--lr", "2e-5",
],
)
# Run the job — Vertex provisions a2-highgpu-1g (A100) machines
model = job.run(
replica_count=1,
machine_type="a2-highgpu-1g", # 1x A100 GPU
accelerator_type="NVIDIA_TESLA_A100",
accelerator_count=1,
)#Cloud TPUs
TPUs (Tensor Processing Units) are Google's custom AI accelerators. They are not GPUs — they are purpose-built for matrix multiplication, the core operation in neural networks.
TPU generations
| TPU Version | TFLOPS (BF16) | HBM | Best For |
|---|---|---|---|
| TPU v4 | 275 | 32 GB | Training medium models |
| TPU v5e | 197 | 16 GB | Cost-efficient inference |
| TPU v5p | 459 | 95 GB | Large model training |
| TPU v6e (Trillium) | 918 | 32 GB | Next-gen training & inference |
#BigQuery ML
BigQuery ML lets you train models directly in SQL. No Python, no infrastructure management.
-- Train a classification model in SQL
CREATE OR REPLACE MODEL `my_project.my_dataset.churn_model`
OPTIONS(
model_type='BOOSTED_TREE_CLASSIFIER',
input_label_cols=['churned'],
max_iterations=50
) AS
SELECT * FROM `my_project.my_dataset.customer_features`
WHERE split = 'train';
-- Predict on new data
SELECT * FROM ML.PREDICT(
MODEL `my_project.my_dataset.churn_model`,
(SELECT * FROM `my_project.my_dataset.customer_features` WHERE split = 'test')
);
#GCP Data Layer for AI
- Cloud Storage: Object storage for training data and model artifacts (GCS buckets).
- BigQuery: Petabyte-scale data warehouse with built-in ML capabilities.
- Dataflow: Stream and batch data processing (Apache Beam).
- Firestore: NoSQL document database for AI application state.
#Head-to-Head: What Each Cloud Does Best
#Real Cost Example: Fine-Tuning a 7B Model
Training a 7B parameter model for 1 epoch on 100GB of data costs approximately:
| Cloud | Instance | Approximate Cost | Notes |
|---|---|---|---|
| AWS | p4d.24xlarge (8x A100) | $200-500 | Most tutorials, largest community |
| GCP | a2-ultragpu-1g (8x A100) | $150-400 | TPU alternative can be 2-3x cheaper |
| Azure | NC96ads_A100_v4 (4x A100) | $180-450 | Enterprise discounts can cut 30-50% |
#How to Choose — Follow the Data
The biggest mistake teams make is choosing a cloud based on a blog post or conference talk. The right cloud is determined by your constraints, not by feature lists. Use this decision tree:
- If your data is in S3 (AWS), the egress cost to move it elsewhere is $0.09/GB. For 50TB, that is $4,500 just to move data. Start with AWS.
- If your data is in BigQuery (GCP), it is already optimized for Vertex AI. Start with GCP.
- If your data is in Azure Blob Storage, Azure ML can access it directly. Start with Azure.
- Need GPT-4 with enterprise security and data residency? Azure is the only option.
- Need Claude, Llama, and Mistral via a single API? AWS Bedrock.
- Need Gemini with tight platform integration or TPU training? GCP.
- A team with 3 years of AWS experience will ship 2-3x faster on AWS than on a "technically superior" platform they have never used. Do not underestimate the cost of context switching.
- A Microsoft Enterprise Agreement can give you 30-50% off Azure. An AWS Savings Plan or GCP Committed Use Discount can similarly reduce costs. Check with your finance team before choosing.
Your company uses Office 365, has 10TB in Azure Blob Storage, and wants to fine-tune GPT-4. Which cloud platform should you choose?
The answer is C. This is a textbook case of "follow the data." Your 10TB is already in Azure Blob Storage (moving it would cost ~$900 in egress fees alone). Your company is already paying for Office 365 (likely has an Enterprise Agreement with Microsoft discounts). And GPT-4 with enterprise security is only available through Azure OpenAI. Every constraint points to Azure. Choosing AWS or GCP here would mean paying more, moving data, and fighting organizational inertia — all for no technical advantage.
#The Grand Comparison
This is the table you will reference every time you pick a cloud for an AI project:
| Dimension | AWS | Azure | GCP |
|---|---|---|---|
| Market share | ~32% | ~23% | ~11% |
| ML platform | SageMaker | Azure ML | Vertex AI |
| LLM service | Bedrock (Claude, Llama, Mistral) | Azure OpenAI (GPT-4, GPT-4o) | Vertex AI (Gemini, Llama) |
| Custom hardware | Trainium/Inferentia | Maia 100 (coming) | TPU v5p/v6e |
| NVIDIA GPUs | p4d (A100), p5 (H100) | NC A100 v4, ND H100 v5 | a2 (A100), a3 (H100) |
| SQL-based ML | Redshift ML | Synapse ML | BigQuery ML |
| Vector search | OpenSearch | AI Search | Vector Search |
| AutoML | SageMaker Autopilot | AutoML in Azure ML | Vertex AutoML |
| Pricing model | Pay-as-you-go, Savings Plans | Pay-as-you-go, Reserved | Pay-as-you-go, CUDs |
| Free tier | 250 hrs/month SageMaker Studio | $200 credit | $300 credit |
#When to Choose Each
Choose AWS when
- You need the broadest service selection
- Your organization is already AWS-heavy
- You want model choice (Bedrock offers Claude, Llama, Mistral)
- You need the most mature ecosystem (most tutorials, most community support)
Choose Azure when
- Your company has a Microsoft Enterprise Agreement
- You specifically need OpenAI models with enterprise security
- You want tight integration with Microsoft 365, Teams, Power BI
- Regulatory requirements demand Microsoft's compliance certifications
Choose GCP when
- You use JAX or TensorFlow and want TPU access
- Data analytics is central to your ML workflow (BigQuery)
- You want Gemini models with tight platform integration
- You value developer experience and clean API design
Tests · Calculate training costs for all three clouds. Verify GCP has the lowest on-demand cost. Add spot pricing and determine the new winner.
#Multi-Cloud and Hybrid Strategies
Most mature AI teams do not use a single cloud. They use multi-cloud strategies:
#Strategy 1: Best-of-Breed
Use each cloud for what it does best:
- GCP for training (TPUs, cost-effective GPUs)
- AWS for serving (broadest deployment options, Lambda for lightweight inference)
- Azure for enterprise LLM access (Azure OpenAI)
#Strategy 2: Primary + Overflow
Pick one primary cloud for 90% of workloads. Use a second cloud for:
- Burst capacity during training spikes
- Access to specific models (Azure OpenAI, Bedrock)
- Disaster recovery
#Strategy 3: Cloud-Agnostic Tooling
Use tools that abstract away the cloud provider:
- Kubernetes (EKS/AKS/GKE) for compute orchestration
- MLflow for experiment tracking (works everywhere)
- Hugging Face for model hosting (provider-agnostic)
- DVC for data versioning (storage-agnostic)
#Common Mistakes
-
Choosing a cloud based on a blog post — Pick based on your organization's existing investments, compliance requirements, and team skills. Not on which cloud had the best marketing this quarter.
-
Ignoring data gravity — If your 500TB data lake lives in S3, moving it to GCS costs $45,000+ in egress fees alone. Start your cloud decision from where your data lives.
-
Using managed services for everything — SageMaker endpoints cost 2-3x more than self-managed inference on EC2. Managed services are worth it early on; self-managed saves money at scale.
-
Forgetting networking costs — Cross-region data transfer, VPC peering, and egress fees add up silently. A model serving system that fetches embeddings from a different region can cost thousands in hidden networking fees.
-
Not negotiating — At scale ($50K+/month), every cloud provider offers custom pricing. Enterprise Discount Programs (AWS), Enterprise Agreements (Azure), and Committed Use Discounts (GCP) can save 30-50%.
#Quick Check
When choosing a cloud provider for an AI workload, what should be the primary deciding factor?
#Key Takeaways
- Data gravity drives cloud choice more than features. If your data already lives in AWS S3, egress fees and migration effort often outweigh any advantage a competing cloud's ML services offer; optimize within your existing ecosystem first
- Each cloud has a distinct AI identity. AWS Bedrock for the broadest managed LLM selection, Azure OpenAI Service for enterprise GPT-4 access, GCP Vertex AI + TPUs for training-intensive JAX/TensorFlow workloads; pick the one that matches your stack
- Cloud-agnostic tooling is the escape hatch. MLflow for tracking, Kubeflow for orchestration, and Docker for packaging keep you portable without sacrificing convenience; build on managed services but abstract the critical interfaces
- Multi-cloud adds capability and complexity in equal measure. Best-of-breed (three clouds for three purposes) requires a dedicated platform team and cross-cloud egress costs; primary + overflow is usually the right balance for teams under 50 engineers
- TPUs beat GPUs on cost-per-FLOP for specific workloads. At scale, TPU v5p can be 2-3× more cost-effective than H100 for large matrix multiplication in JAX; but PyTorch workloads belong on NVIDIA unless you can rewrite them