跳到正文
原文
Artificial Intelligence News(网页)· Artificial Intelligence News(网页)·· 6 天前AI 评分52

2026 年生成式 AI 开发:面向生产环境的系统构建

AI 导读

这篇分析指出,2026 年生成式 AI 开发的重心已从单一 LLM 转向模型周边系统,包括按任务路由多模型、上下文工程、通过 MCP 接入工具的智能体、以及安全治理。

正文

In 2024, much of generative AI development focused on getting one basic loop to work: a user entered a prompt, a large language model generated a response, and the result appeared on the screen. Making that loop reliable enough for a real product still requires serious engineering work. 

By 2026, a developer can connect to a frontier model simply through an API, and so can every competitor. So, the hard part has moved elsewhere. The months of work now go into the system around the model: the context it receives, the tools it can use, the checks on its output, the permissions that constrain it, and the logs that explain what happened afterward. 

Figure 1. The request path in 2024 and in 2026. Illustration by the author. 

This article looks at generative AI development as it stands in September 2026. For companies investing in generative AI development services, the model has become only one component of the product. Most of the engineering effort now goes into workflows, integrations, evaluation, security and the ability to switch models without a rewrite. 

What changed in generative AI development since 2024 

The shift is easiest to see side by side. 

Area Typical in 2024 Typical in 2026 
Interface Chat window Agents embedded in business workflows 
Models One LLM Several models, routed per task 
Core skill Prompt engineering Context engineering 
Knowledge RAG as the default answer RAG plus tools, memory and structured data 
What the AI does Answers questions Carries out tasks within set limits 
Quality measure Model accuracy End-to-end task success 
Cost measure Price per million tokens Cost per completed business task 
Deliverable Prototype or pilot Production software system 

The model is becoming a replaceable component 

AI models are improving incredibly quickly. The best option today may not be the best one six months from now. There are plenty of reasons for that: prices change, models get retired, providers have outages, and cheaper alternatives can suddenly become good enough. 

That makes it risky to build a product around a single model. For instance, Bloomberg reported that the legal AI company Harvey  saw its gross margin fall from about 50% at the start of 2026 to roughly minus 50% by June, after its token usage increased twentyfold. In August, Harvey released its own model, post-trained on Moonshot’s open-weight Kimi K3, and its margins turned positive again. It shows that your application should be able to switch between models when cost, performance or business requirements change. 

That is what a routing layer does. Instead of sending every request to the same model, it chooses the most suitable one based on factors such as task complexity, speed, context length, data sensitivity, price and modality. 

In a customer operations product, that split might look like this: 

Task Model choice Why 
Classifying incoming support tickets Small, fast model High volume, simple labels 
Drafting a reply to a billing dispute Frontier reasoning model Policy nuance and customer impact 
Extracting fields from scanned invoices Vision and document model Layouts, stamps, handwriting 
Searching internal code that contains client data Open-weight model on own infrastructure Data stays in-house 

Design the product around capabilities such as classification, extraction and multi-step reasoning, and let the router map each one to whichever model handles it best this quarter. 

Context engineering is now the hardest part 

Context engineering is the work of deciding what information the model receives on each call. It has overtaken prompt wording as the main challenge, since a well-phrased prompt cannot make up for missing or wrong data. This is why generative AI development services increasingly focus on data access, context selection and system architecture rather than prompt writing alone. 

Take a simple question: “Why was I charged twice?” A good answer depends on the customer’s subscription history, their last few invoices, the payment records and the refund policy, and some of that should stay hidden from this particular customer. 

Each of those pieces lives somewhere different. The refund policy and help articles are text, so  AI systems handle them well: the system searches the documents and pastes the relevant passages into the prompt. Invoices and payments are rows in a database, and document search won’t find them. For those, the model needs controlled access through an API or a semantic layer, which pins down what terms like “billing period” mean so that every query uses the same definition. 

Memory is a third source. Session memory holds the current conversation, user memory keeps facts across conversations, and application state tracks where a business process stands. Keeping the three apart stops a system from repeating outdated promises or acting on stale status. 

The final decision is how much to include. Large context windows make it tempting to send everything. Each additional token adds cost and latency, though, and irrelevant material gives the model more to misread, so good context engineering selects, ranks and summarizes before each call. 

Once a system reliably has the right information, the next step is to let it act on that information. 

From answering to acting 

Another major shift is from generative AI to agentic AI. A traditional chatbot follows a relatively simple pattern: User asks something → model generates an answer. 

An agent can operate differently: 

User defines a goal → agent decides what needs to happen → agent chooses tools → agent performs actions → agent checks the result → agent continues or returns the outcome. 

For this part, consider a customer support system. Instead of explaining how a customer can change a subscription, an agent could check the account, confirm whether the requested change is allowed, calculate the new price, update the subscription, record the action in the CRM, and send confirmation. 

That is much more useful. It is also much more difficult to build safely. The application needs connections to external systems and clear rules around what the agent can do. 

Standards are making some of this integration easier. The Model Context Protocol, or MCP, standardizes how agents connect to tools, APIs, databases, and other resources. Agent2Agent, or A2A, addresses another problem: communication between independent agents. In simple terms, MCP helps an agent use tools, while A2A helps agents work with other agents. 

This is an important change for generative AI software development services. The work increasingly involves designing workflows and tool access rather than simply wrapping an interface around an LLM. 

Faster coding moves the bottleneck 

AI coding agents are now part of everyday engineering. The JetBrains Developer Ecosystem Survey of more than 15,000 professional developers found that between May and July 2026, 90% used AI coding agents at work at least weekly and 68% used them daily. 

Faster code production does not automatically mean faster delivery. Business Insider reported on September 18 that about 80% of Oracle employees adopted ChatGPT Enterprise and Codex within three months of the spring rollout, and some work that used to take a team two to three quarters now takes about a week. Releases to customers did not accelerate at the same pace, and co-CEO Clay Magouyrk told staff that testing, validation, deployment and release management have to be redesigned. 

The reason is that software delivery is a chain of steps, and speeding up one step piles work onto the next. Sombra saw the same effect in a published case: six AI agents covering analysis, design, development, review, QA and project management built a lab resource portal for an energy storage company in seven days, against an estimated three months for a conventional team, with Sombra engineers handling only exceptions. At that speed, the clarity of requirements and the time available for review set the pace. 

For companies buying generative AI software development services, investment should therefore follow the bottleneck into architecture, automated testing, evaluation and release engineering. Of these, testing changes the most when the software is generative. 

Cost per completed task is the number that matters 

Token prices kept falling through 2026, yet many companies saw their total AI bills rise. The Financial Times reported in June that Amazon, Walmart, Cisco, Uber and Meta had introduced spending caps, discouraged wasteful use or pushed staff toward cheaper models. Workato’s chief information officer, Carter Busse, told the FT that the company’s spend rose sevenfold after Anthropic switched to token-based pricing in May: “We created a monster.” 

Agents explain much of the gap. A single task can involve dozens of model calls for planning, tool use, retries after errors, rereading the conversation history and checking the output. A low price per token multiplied across a long loop can still produce an expensive task, which is why the useful measure is cost per completed task. 

The illustrative example below compares three designs for a support workflow that makes 12 model calls per ticket, with unresolved tickets passed to a human at $6 each: 

Design AI cost per ticket Resolved by AI Total cost per ticket 
Frontier model on every step $0.72 78% $2.04 
Small model on every step $0.05 41% $3.59 
Small model first, frontier model on hard steps $0.22 75% $1.72 

Illustrative figures: $0.06 per frontier model call, $0.004 per small model call, $6 per ticket handled by a person. 

The cheapest model per token produces the most expensive outcome, because it resolves too few tickets and passes the rest to people. The routed design described earlier comes out cheapest, since it pays for a frontier model only where the decision is hard. Caching repeated context, trimming conversation history and setting step and spend limits for each agent run reduce costs further. 

Spending limits are one way of controlling what an agent can do. Security needs several more. 

Security and governance for agents that act 

Once a model can call tools, its mistakes turn into actions: a refund issued, a record deleted, an email sent to the wrong person. 

Security 

The most common attack is prompt injection, in which instructions are hidden in content the agent reads, such as an email or a web page. OWASP’s State of Agentic AI Security and Governance report, updated in June 2026, links prompt injection to six of the ten categories in its Top 10 for Agentic Applications. The report flags one combination as especially risky: an agent that has access to private data, reads untrusted content and can send data out. Many useful business agents have all three, so the controls have to be part of the design: 

  • dedicated credentials for each agent, limited to what it needs 
  • separate tools for reading and writing, with write access granted narrowly 
  • human approval for irreversible or high-value actions 
  • all retrieved and user-supplied content handled as untrusted input 
  • an audit log entry for every tool call 

Governance 

Regulators now require some of the same controls. Parts of the EU AI Act’s transparency rules have applied since August 2, 2026. Under Article 50, people must be told when they are interacting with an AI system unless it is obvious, and AI-generated content must carry machine-readable marking. Systems already on the market before that date have until December 2, 2026 to meet the marking requirement, and the Digital Omnibus moved most high-risk obligations to December 2027 or later. Companies in the USA also face a growing number of state AI laws. 

In engineering terms, compliance means an AI disclosure in the interface, content marking where it applies, and records showing which model, prompt and data produced each output. These requirements shape the data model and the logging design, so they belong in the first architecture review. 

A reference architecture for 2026 

The layers described so far fit together in one architecture. 

Figure 2. A reference architecture for production generative AI. Illustration by the author. 

The application layer handles identity, permissions and AI disclosure. Below it, an orchestration layer runs the agent loop, enforces budgets and asks for human approval when an action crosses a threshold. It draws on three resources: context from documents, data and memory; tools reached through MCP; and models selected by a router. Evaluation and guardrails check every output, while observability, cost tracking, security and versioning span the whole system. 

Build, buy or combine 

Seen this way, the build-or-buy question becomes a question about layers. Companies usually buy the commodity layer: foundation models, embeddings, OCR, speech recognition and hosting. They configure the platform layer, where CRMs, help desks and cloud providers now include agent features and off-the-shelf generative AI development solutions cover common cases. And they build the differentiation layer: their own workflows, proprietary data, domain logic, integrations, evaluation sets and approval rules, which competitors cannot copy by using the same APIs. 

That last layer is where custom generative AI development services add the most value. When comparing a generative AI development company, look for evaluation results from production systems, cost per completed task and a clear plan for switching models as requirements change. Sombra, for example, supports this process from consulting and data preparation through proof of concept, MVP and full application development. 

Summary 

Generative AI development in 2026 is mostly systems engineering. Capable models are available to every company, and they are also the part of the stack most likely to change within a year. The work that decides whether a product succeeds sits around them: 

  • routing that keeps models replaceable 
  • context engineering that gives each call the right data 
  • tool access through MCP, with tight permissions 
  • evaluation based on completed tasks 
  • cost measured per outcome 
  • security and compliance designed into the architecture 

To start, pick one workflow, define a good outcome, build a thin end-to-end version on real data, create the evaluation set early and add autonomy in stages as the results justify it. 

来源:Artificial Intelligence News(网页) · artificialintelligence-news.com