Artificial Intelligence

Safely Releasing Frontier Models to Customers

Safely Releasing Frontier Models to Customers

It’s our goal for AWS to be the most secure place to run any workload, and in support of that we’ve been deeply investing in security across our services since AWS’s inception more than two decades ago. Our AI services like Amazon Bedrock are built on this foundation and with the same focus. 

Generate images and video with vLLM-Omni on SageMaker AI - Part 2

Generate images and video with vLLM-Omni on SageMaker AI – Part 2

Deploy two generative media models from one AWS vLLM-Omni Deep Learning Container on Amazon SageMaker AI. Generate an image with FLUX.2-klein through real-time inference, then animate it into video with Wan2.1-VACE through asynchronous inference, and retrieve the MP4 from Amazon S3.

Automating Amazon Textract adapter lifecycle management across accounts

Automating Amazon Textract adapter lifecycle management across accounts

Learn how to operationalize Amazon Textract Custom Queries adapters for production: infrastructure as code with AWS CloudFormation and Terraform, a cross-account adapter promotion process, a pre-classification routing pattern for multiple form versions, and production security controls such as VPC endpoints, encryption, and least-privilege IAM.

Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput

Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput

Learn how to scale Mixture-of-Experts (MoE) reinforcement learning on Amazon EKS using Elastic Fabric Adapter (EFA) and DeepEP. This post presents an architecture that combines Amazon EKS, EFA, and Amazon S3 and increased aggregate reinforcement learning rollout throughput by 40% for large-scale RLHF and GRPO training.

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod

Learn how to run SkyRL, an open-source reinforcement learning framework, on Amazon SageMaker HyperPod to post-train a Qwen3-VL-8B vision-language model with GRPO. This walkthrough covers building the container image, launching a Ray cluster from SageMaker Studio, submitting and monitoring the job, and hosting the trained LoRA adapter for inference.

NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

NarrateAI: production-ready LLM quality assurance on Amazon Bedrock

NarrateAI delivers production-ready LLM quality assurance on Amazon Bedrock. This post details five techniques—adaptive pipeline orchestration, cross-account multi-model failover, real-time streaming evaluation, composite evaluation, and data accuracy verification—that reach about 99% numerical accuracy while streaming responses in real time.

How Datacor built self-service rental analytics with Amazon Quick Sight

How Datacor built self-service rental analytics with Amazon Quick Sight

Learn how Datacor built a self-service rental analytics experience for gas and welding distributors by embedding Amazon Quick Sight dashboards and natural language querying into its TrackAbout platform, powered by an automated cross-cloud data pipeline and multi-tenant row-level security.

Multi-Region training with Amazon SageMaker HyperPod and Qumulo

Multi-Region training with Amazon SageMaker HyperPod and Qumulo

Amazon SageMaker HyperPod and Cloud Native Qumulo let you place training compute in one AWS Region while keeping your dataset in another. This post shares the architecture and validation results from a cross-Region training run, where a remote cluster matched a co-located cluster’s throughput after a brief NeuralCache warmup.