IBM releases Granite Time Series PatchTST-FM-r2, a 385M-parameter open zero-shot forecasting model ranking #2 on GIFT-Eval
IBM released Granite Time Series PatchTST-FM-r2, a ~385M-parameter time-series foundation model (successor to PatchTST-FM-r1) with a redesigned conformer-based architecture (self-attention + temporal convolution), context length up to 8,192, probabilistic forecasting via 99-quantile prediction head, support for missing-value imputation, and dual Apache-2.0/OpenMDW-1.0 licensing. Model weights, architecture, inference pipeline, and reproduction code were published on Hugging Face.
Entities: IBM, IBM Research, Granite Time Series PatchTST-FM-r2, PatchTST-FM-r1, GIFT-Eval, Hugging Face
1 primary
What happened
IBM released Granite Time Series PatchTST-FM-r2, a roughly 385-million-parameter open-weight time-series forecasting model, as the successor to PatchTST-FM-r1. It has a redesigned architecture combining self-attention and temporal convolution, handles context lengths up to 8,192, produces probabilistic forecasts across 99 quantiles, and can handle missing data. Weights, code and a reproduction pipeline are published on Hugging Face under dual Apache-2.0/OpenMDW-1.0 licences, and IBM reports it ranks #2 on the GIFT-Eval leaderboard among zero-shot models.
Why it matters
This gives developers and enterprises a single pre-trained model that can forecast on new datasets without retraining, and the permissive commercial licence removes a real barrier that restrictive research licences often impose. It matters most to teams doing demand forecasting, capacity planning or anomaly detection who want to avoid building bespoke per-dataset models. The practical upside is real but narrow: this is one component in a forecasting pipeline, not a general capability shift, and adoption depends on integration effort and how it performs on a given company's actual data rather than benchmark data.
What is noise
The "SOTA" framing in the headline is true only within a carefully carved-out subset (permissive-licence, replicable, zero-shot models), and the model actually ranks #2 overall, not #1. It's an incremental r1-to-r2 update in an already crowded field (TimesFM-3, Chronos-2, Timer-S1, Toto all compete here), and the Confluent integration mention reads as a partnership plug rather than independent validation. Coverage should be read as a solid technical release, not a breakthrough.
Watch next
- 01Whether independent users replicate the claimed GIFT-Eval metrics (CRPS 0.467, MASE 0.6846) outside IBM's own benchmarking.
- 02Download and usage figures on Hugging Face over the next 1-3 months compared to Chronos-2 and TimesFM-3 to gauge real adoption versus PR interest.
- 03Whether IBM or third parties publish head-to-head comparisons on real enterprise datasets (not just GIFT-Eval) showing performance against fine-tuned, per-dataset models.
Evidence
1 linkedCoverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Google DeepMind launches AlphaGenome Atlas, a free public database of predicted effects for 9 billion possible human genome variants8 Sept 202679
- Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval3 Sept 202679