Cortex Agent and Cortex Analyst evaluations: Default metric version and judge model changes (Pending)¶
Snowflake is changing evaluation defaults because the v1 judge model, claude-4-sonnet, reaches end-of-life on October 14, 2026. It has been in the legacy state since August 12, 2026, so accounts that didn’t use it before that date already can’t run v1 evaluations. The new judge models also have context windows of about 1 million tokens, which reduces failures on long agent traces.
Important
This change isn’t part of a behavior change bundle. It is expected to be enabled by default starting October 13, 2026. Dates are subject to change.
Behavior change¶
This change affects the following metrics when they omit the setting shown or set it to auto:
- Cortex Agent system metrics (
answer_correctness,logical_consistency,tool_selection_accuracy, andtool_execution_accuracy) and the Cortex Analystsql_correctnessmetric:version - Cortex Agent custom metrics:
model
- Before the change:
System metrics default to
v1, and custom metrics default toclaude-4-sonnet.- After the change:
System metrics default to
v3, and custom metrics default toclaude-sonnet-4-6. For both, Snowflake usesopenai-gpt-5.4ifclaude-sonnet-4-6isn’t allowed or available for your account.
Explicitly specified versions and models don’t change.
Impact and preparation¶
Scores from different metric versions aren’t directly comparable. Scores for metrics that use a judge model can change. tool_selection_accuracy doesn’t use a judge model, so its scores don’t change. Evaluation costs can also change because judge models have different credit rates; track them with the CORTEX_REST_API_USAGE_HISTORY view.
No action is required to adopt the new defaults. Before rollout:
- Test system metrics with
version: "v3", and recalibrate thresholds in dashboards, alerts, and CI/CD checks. - To control when you upgrade, set a system metric
version, such asv2, or a custom metricmodel. Metrics pinned tov1orclaude-4-sonnetstop working at end-of-life. - Confirm that
claude-sonnet-4-6oropenai-gpt-5.4is allowed for your account and the role that runs the evaluation, and is available through your cross-region inference configuration. See judge models and cross-region inference.
See also¶
- Cortex Agent evaluation metric versions
- Cortex Analyst evaluation metric versions
- Cortex Agent custom metrics
- Cortex model deprecations for August 2026
- Behavior change policy
Ref: 2442