蚂蚁集团开源 AKernel:面向 AI 智能体的可编程数据中心基础设施
蚂蚁集团开源 AKernel(Agent Kernel),一个将整个数据中心视为 AI 智能体可编程扩展的分布式内核。它通过统一 Python SDK 实现计算、网络和存储的编程控制,支持单机、私有 K8s 及多云(阿里云、华为云)一键部署,冷启动仅需 40 毫秒。AKernel 代码由 AI 大量贡献,并内置基于 OpenTelemetry 的全栈可观测性。
蚂蚁把数据中心变成 Agent 的可编程扩展,对于需要大规模弹性编排 Agent 的团队来说,比用 Kubernetes 手写 YAML 直接得多,不过目前严重依赖阿里云和华为云,通用性还差一截。
Overview
AKernel (Agent Kernel) is a distributed kernel that combines the performance of AFaaS with the architecture of openYuanrong, enabling true "datacenter use" — treating the entire datacenter as a programmable extension of your AI Agent.
Traditional infrastructure tools (IaC, Kubernetes-native platforms, multi-cloud Terraform, and vendor-specific CDKs) are designed for provisioning infrastructure, not operating it. They fall short for AI agents, RL training, and data pipelines that require runtime elasticity, dynamic workflows, and programmatic access to datacenter capabilities.
AKernel solves this with five key advantages:
Datacenter Use: Unified Programming Interface
A single Python SDK (akernel_sdk) to programmatically control compute, networking, and storage — no YAML, no manual orchestration. Multi-language support is planned.
from akernel_sdk import Sandbox with Sandbox(cpu=2000, memory=4096) as sb: result = sb.commands.run("echo 'hello from AKernel'") print(result.stdout)
One-Click Deployment: From Laptop to Multi-Cloud
One all-in-one image, multiple deployment targets — deploy in under 10 minutes:
| Mode | Target | Scale | Deploy Time |
|---|---|---|---|
| Standalone | Single machine | 1 node | ~1 min |
| Private K8s | On-premise cluster | 100 nodes | ~5 min |
| Multi-Cloud | Alibaba Cloud, Huawei Cloud, etc. | 100 nodes | ~10 min |
Secure Isolation with Extreme Performance
- 40 ms cold start*: Fork-based launch with lazy loading for near-zero startup latency
- Sandbox isolation: gVisor by default, with Kata Containers and Firecracker available on KVM-capable nodes
- Same-node recovery: Checkpoint runsc and Firecracker workloads and reload the same logical sandbox
* Planned for an open-source release and not available in AKernel v0.1.0.
AI-Native Development and Operations
- AI-contributed codebase: Significant portions of AKernel code are authored by AI, enabling rapid iteration
- AI-driven operations: Cluster deployment, day-2 operations, and troubleshooting powered by AI agents
- Agent-friendly: Built as infrastructure that AI agents can programmatically control and reason about
Full-Stack Observability
Built-in OpenTelemetry (OTEL) integration provides complete observability out of the box — not just resource management, but the full picture:
- Metrics: Prometheus-based collection for compute, networking, and sandbox performance
- Dashboards: Pre-configured Grafana dashboards for real-time cluster monitoring
- Tracing & Logging: End-to-end request tracing and centralized log aggregation
Quick Start
Quick Navigation
- 💡 Examples - AKernel SDK examples and use cases
- 🏗️ Architecture - System design and components
- 🚀 Deployment - Installation and configuration guide
Bootstrap a Cluster
AKernel provides guided Terraform deployment for Alibaba Cloud ACK and Huawei Cloud CCE. Clone the repository and prepare the cloud credentials before selecting one of the image options below.
The workflow requires Terraform, Docker, Helm, kubectl, Python 3, GNU Make, and cloud credentials with permission to create the required infrastructure. For Alibaba Cloud, export the credentials and region first:
git clone --recurse-submodules https://github.com/akernel-dev/akernel.git cd akernel export ALICLOUD_ACCESS_KEY="<your-access-key-id>" export ALICLOUD_SECRET_KEY="<your-access-key-secret>" export ALICLOUD_REGION="cn-hangzhou"
Option 1: Use the Official Image
Use the public AKernel image from Docker Hub:
make config VENDOR=aliyun \ IMAGE_REPOSITORY=akerneldev/all-in-one \ IMAGE_TAG=latest make deploy
Option 2: Build from Source
Configure a registry that your cluster can access, then build and push the all-in-one image before deployment:
make config VENDOR=aliyun \ IMAGE_REPOSITORY=registry.example.com/akernel/all-in-one \ IMAGE_TAG=your-release-tag docker login registry.example.com make build make push make deploy
See the Deployment Guide for prerequisites, cloud-specific configuration, deployment verification, and cluster cleanup, and the Build Guide for development details.
Create a Sandbox
Install the Python SDK. The default installation includes the
openyuanrong-sandbox backend:
# PyPI python -m pip install akernel-sdk # Source python -m pip install ./sdk/python # Also install the deprecated actor compatibility backend python -m pip install "akernel-sdk[openyuanrong-sdk]"
The actor-based openyuanrong-sdk backend is deprecated and retained only for
compatibility with existing applications. New applications should use the
default openyuanrong-sandbox backend. When the actor extra is installed,
both backend packages are present and openyuanrong-sandbox remains the
automatic default. Set
AKERNEL_BACKEND=openyuanrong-sdk before importing akernel_sdk to select
the actor backend:
export AKERNEL_BACKEND=openyuanrong-sdkWhen openyuanrong-sdk is used from a YuanRong function, the SDK process
inherits runtime paths configured by builder/scripts/entryfile.sh. That
entrypoint exports PYTHONPATH and, in some runtime layouts,
LD_LIBRARY_PATH; both variables are inherited by the application and its
child processes. PYTHONPATH prepends the runtime site-packages directory and
can change import resolution or shadow application dependencies.
LD_LIBRARY_PATH prepends runtime library directories and can change native
library resolution, causing ABI or version conflicts.
Configure the AKernel environment:
export AKERNEL_SERVER_ADDRESS="<your-akernel-server-address>" export AKERNEL_TOKEN="<your-akernel-token>"
Use the SDK to create and interact with a sandbox:
from akernel_sdk import Sandbox with Sandbox(cpu=1000, memory=2048) as sandbox: result = sandbox.commands.run("echo 'hello from AKernel'") print(result.stdout) sandbox.files.write("/tmp/hello.txt", "hello from the SDK") print(sandbox.files.read("/tmp/hello.txt"))
Experimental gVisor sandboxes can request an exact NVIDIA GPU model:
with Sandbox(xpu="gpu:l20:1") as sandbox: print(sandbox.commands.run("nvidia-smi -L").stdout)
GPU sandboxes require a compatible NVIDIA node and the gVisor runsc
runtime. storage_mb is measured in MiB and is supported by runsc and
Firecracker.
See the complete basic usage example, the sandbox runtime example, and the other SDK examples for more operations.
Architecture
System Components
Node-Level Infrastructure
- Sandbox runtimes: gVisor by default; Kata Containers and Firecracker on KVM-capable nodes; and an explicitly enabled native Linux runc backend
- sandboxd: Sandbox lifecycle daemon with pluggable sandbox runtime integration
- distill-fs: Rust-based FUSE filesystem for lazy rootfs access, chunk caching, and deduplication; packaged from a static GitHub Release with its version and checksum pinned in AKernel
Cluster-Wide Services
- Distributed Scheduler: Workload-aware placement and scaling
- API Gateway: Unified interface for all operations
- Object Storage: Raw and Nydus rootfs images and read-only sandbox mounts
- Cloud Provisioning: Terraform modules for Alibaba Cloud ACK and Huawei Cloud CCE
How It Works
- Agent Submits Workload: Through unified API or SDK
- Scheduler Places Sandbox: Selects a worker based on requested CPU, memory, storage, accelerator model, and available capacity
- Sandbox Created: Prepares the rootfs and network and starts the selected sandbox runtime on the worker
- Workload Executes: In secure, isolated sandboxes
- Resources Recycled: Deletes the sandbox and returns its capacity to the cluster
Roadmap
- Kata Containers runtime on KVM-capable nodes
- Firecracker microVM runtime on KVM-capable nodes
- Optional native Linux runc runtime
- Stateful sandbox network ACLs for CIDRs, domains, protocols, and ports
- Fork-based sandbox launch based on gVisor
- Same-node checkpoint recovery for runsc and Firecracker
- Support for GKE and AWS
- Cgroup v2 node support
License
AKernel is licensed under the Apache License 2.0.
来源:蚂蚁 inclusionAI:GitHub 新仓库 · github.com