live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
analysisUNDERREPORTED

Salesforce’s Moirai staffing gains need an operational forecast test card

The vendor reports double-digit holiday staffing improvements, but practitioners still need the baseline, error definition and intervention path before treating them as transferable evidence.

Forecast metrics versus staffing outcomes.
AI-generated illustration
By The News Desk· Sep 11, 2026the quick take — two AI hosts go live when you do

Salesforce has attached an unusually concrete business outcome to its Moirai time-series model: customer feedback showed a 14% improvement in understaffing accuracy and a nearly 40% improvement in overstaffing accuracy during the holiday period. The company’s Sept. 3 video page also says Moirai has passed 30 million Hugging Face downloads and ranks highly on several time-series benchmarks. (Salesforce News)

Those numbers make the page worth more than a routine model reminder. They also illustrate why an enterprise forecasting claim needs two scorecards: one for the model and another for the decision process built around it.

The model evidence and the staffing evidence are different

The original Moirai paper describes a universal forecasting transformer trained on the LOTSA archive, which contains more than 27 billion observations across nine domains. It reports competitive or superior zero-shot forecasting performance against full-shot models and links the public code, data and weights. (Moirai paper)

A later Salesforce model card says Moirai 1.1-R improved normalized mean absolute error by roughly 20% on low-frequency yearly and quarterly cases across 40 Monash datasets. The same card labels the release research-only and recommends downstream evaluation before deployment, especially in high-risk settings. (Moirai 1.1-R model card)

Neither result, by itself, explains the staffing gains. The Salesforce video page does not identify the customer, sample size, forecast horizon, baseline system, error formula or whether “accuracy” measures demand forecasts, schedules, or the final under/overstaffing outcomes. It also does not say whether the 14% and nearly 40% figures are relative or absolute improvements. (Salesforce News)

What a buyer should ask for

A usable test card would separate four layers:

  1. Forecast quality: error by horizon, season and location against the incumbent method.
  2. Decision policy: how a forecast becomes a staffing recommendation, including constraints and override rules.
  3. Operational outcome: understaffed and overstaffed intervals, service levels and labor cost.
  4. Intervention record: how often planners reject or modify the recommendation, and why.

That separation matters because a model can reduce aggregate forecast error while still missing the peaks that drive queues, overtime or idle capacity. Conversely, a modest forecasting improvement can create a larger business gain if the scheduling policy is sensitive to the corrected cases.

Salesforce’s reported holiday result is therefore a credible deployment lead, not yet a portable benchmark. The next useful disclosure is not another leaderboard rank. It is the denominator and evaluation design connecting Moirai’s forecast to the staffing decision.

Filed by The News Desk. Corrections: desk@nl2sql.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.