Snowflake is introducing dynamic model routing across its Cortex AI platform as enterprises look for ways to control the growing cost and complexity of running artificial intelligence workloads.

The new capability, announced August 18, allows Cortex AI Gateway to automatically select models based on factors including quality, cost and workload requirements. Instead of sending every request to the same large model, Snowflake can route simpler tasks to less expensive models while reserving more capable frontier models for workloads requiring deeper reasoning.

The approach is designed to address a problem that is becoming more significant as companies move AI agents and applications from pilots into production: the most powerful model is not necessarily the most economical choice for every task.

Automating Model Selection

Dynamic model routing evaluates individual requests and determines which available model provides an appropriate balance between performance and cost.

Lower-complexity and repetitive tasks can be sent to more efficient models, while requests requiring more sophisticated reasoning can be directed to frontier models.

The capability is being integrated with Snowflake CoCo and Snowflake CoWork and can also be used by third-party AI agents connected through Cortex AI Gateway.

Administrators retain control over which models and providers users can access, which could be particularly important for organizations dealing with compliance requirements or regional restrictions.

Snowflake says routing decisions can also change as model pricing and performance evolve without requiring customers to rebuild their AI applications.

Snowflake Targets AI Token Efficiency

Snowflake’s internal testing suggests the approach could generate meaningful efficiency improvements.

In one evaluation, an AI agent using dynamic model routing built a dbt pipeline with up to three times greater token efficiency compared with an approach relying exclusively on a frontier model while maintaining the same quality.

Another internal test found engineering teams completed the same number of pull requests with 25% greater token efficiency. Snowflake notes that results can vary depending on workload and configuration.

The company is positioning those improvements around what it calls “intelligence efficiency,” or the ability to translate spending on models, compute, data and context into useful business outcomes.

Open Models Expand the Routing Pool

Snowflake is also expanding the models available through Cortex AI.

The company plans to add DeepSeek-V4-Flash 0731 and GLM-5.3 alongside models already available from providers including Anthropic, OpenAI, Google, Meta and Mistral.

A larger model pool gives the routing system more options when deciding how to balance cost and performance.

Snowflake said its AI Research Team found DeepSeek-V4-Flash scored 74.4% on data engineering tasks in its testing, while GLM-5.2 scored 62.8% and used fewer tokens than any other model evaluated.

More Control Over Enterprise AI Spending

Snowflake is pairing automated routing with additional controls designed to help organizations track and manage AI consumption.

Cortex AI Gateway provides administrators with visibility into token usage and costs, while Snowflake CoCo can use existing role-based access and tagging capabilities to establish default models, assign usage to teams or cost centers and create per-user quotas.

Organizations can also receive notifications when AI consumption approaches defined spending limits.

As enterprises deploy larger numbers of AI agents, those controls could become increasingly important. AI infrastructure management is shifting from simply providing access to models toward continuously deciding which models should handle particular workloads and how much organizations are willing to spend on them.

Dynamic routing gives Snowflake a way to automate part of that decision-making process while keeping model access and enterprise data governance within the same platform.

Leave a Reply

Your email address will not be published. Required fields are marked *

Latest News