1. What are the authors trying to do (no jargon)?
A: The authors aim to make time series forecasting (predicting future values from past data) better when there are multiple channels (variables). They propose a method to group similar channels together to improve prediction accuracy and help the model generalize to unseen data. The approach balances treating each channel individually to maintain accuracy and using relationships between channels to support generalization.2. How was it done before their work, and what were the limits of prior work?
A: Prior work handled multiple channels in two main ways:- Channel-Independent (CI): Train a separate forecasting model per channel. This often improves accuracy on each individual series. However, it ignores any relationships between channels. It also fails to generalize well to unseen time series.
- Train one model on all channels jointly. This captures interactions (e.g. correlation between temperature and humidity), but it often "oversmooths" the data. The model's output becomes too averaged across channels. CD models also struggle if channels are very different (low similarity) because one set of parameters tries to fit all series.
3. What is new in their approach?
A: The authors introduce the Channel Clustering Module (CCM), which:- Groups channels based on their similarities, rather than treating them all separately or together.
- Uses "cluster identity" instead of individual channel identity to manage groups while allowing interactions within them.
- Learns cluster prototype embeddings using cross-attention to capture representative time series patterns for each cluster.
- Enables zero-shot forecasting by assigning unseen channels to the closest learned cluster prototypes at test time without retraining.
- Is model-agnostic and can be added to different forecasting architectures like transformers, linear models, or CNNs.
4. What impact does the new approach have?
A: CCM consistently improves forecasting accuracy across diverse models and datasets. It enables zero-shot forecasting by assigning unseen time series to learned cluster representations. Finally, CCM improves interpretability by revealing similarity structures among channels.5. How do the authors evaluate their proposal?
A: CCM was tested on real-world datasets and compared with past methods:- For long-term forecasting, the authors used 9 datasets (like ETTh1, Weather) with horizons up to 720 steps. CCM improved performance in 90.27% of MSE cases and 84.03% of MAE cases.
- For short-term, the authors tested on M4 and a stock dataset and outperformed models like DLinear by 11.62% (on M4).
- For zero-shot forecasting CCM was tested on ETT datasets and showed improvements in all cases, with up to 11.13% better for PatchTST.
- Compared CCM with a regularization method (PRReg), beating it in most cases.
This section explains why specific models and datasets were chosen for this project.
Chose the Beijing Multi-Site Air-Quality Dataset.
- Many datasets used in the paper involve long-horizon multivariate forecasting (predicting multiple variables), so selected as similar dataset.
- The dataset includes hourly air pollutants data from 12 air quality monitoring sites in Beijing. One can perform zero-shot learning on this dataset by training on one region and testing on another.
Chose the DLinear model which is originally a Channel Independent (CI) architecture.
- Had access to the free tier of Google Colab which offers ~3.5 hours of runtime. Opted for a model with fewer parameters as it is expected to keep training time low and converge faster.
- The paper shows a strong performance improvement for DLinear on long-term forecasting tasks.
- CCM also demonstrates stronger improvements on originally CI-based models in the zero-shot setting.
The files in the src/ directory are a minimal rewrite of the original TimeSeriesCCM codebase, but for only the DLinear model. The following files are included (See Code Fixes for more details):
main_dlinear.py: Script for training and evaluating the DLinear model. This is a modified version ofmain.py.data/data_loader.py: Class for loading MTS data.
models/Dlinear.py: The DLinear model implementation.attention.py: Same code as in the original implementation.layers.py: Only includes the cluster assigner definition for the DLinear model.patch_layer.py: Only includes the cluster-wise linear definition for the DLinear model.
exp/exp_basic.py: The base experiment class.exp_ccm.py: The experiment class for the DLinear model.
utils/: Same code as in the original implementation.
NOTE: The notebooks used for execution contain the code from src/ pasted into a single file. On VMs, I would have run the code using the src/ directory structure.
I addressed some issues in the code, with minor assistance from this GitHub issue.
-
main_dlinear.py: A modified version ofmain.pyfor the DLinear model. Reorganized the arguments into a cleaner format. Fixed default values for some experiment parameters: for example, updated the CUDA device fromcuda:2tocuda:0, and changed the experiment class fromExp_TVtoExp_CCM. Removed wandb logging as it can be tricky to set up and use within Google Colab. -
models/Dlinear.py: Updated line 56 to usebatch_sizeas the number of channels for univariate datasets (like M4, stock); otherwise,args.data_dim. -
models/layers.py: The original code returnscluster_embwith shape[batch_size, n_cluster, d_model]. The downstream code uses this:prob = torch.mm(self.l2norm(x_emb), self.l2norm(cluster_emb).t()).reshape(bs, n_vars, self.n_cluster)where
cluster_embshould be[n_cluster, d_model]. To fix the shape mismatch, added a mean overdim=0(the batch dimension), so the final shape becomes[n_cluster, d_model]. -
exp/exp_ccm.py: As described in the paper, the cluster loss relies on the cluster membership matrix M and the similarity matrix between channels, S:However, the original
get_similarity_matrixmethod computes similarity between batch samples, not channels. To address this, added a new method calledchannel_similarity_matrixthat correctly computes channel-wise similarity. Partially implemented using above GitHub issue.
The datasets and model checkpoints may be accessed here.
The experiment runs may be accessed at notebooks/.
The training dataset uses data from two stations: Aotizhongxin and Changping. The zero-shot dataset uses the opposite combination for each respective training dataset. Linear interpolation was used to fill in missing values instead of averaging. This is because interpolation preserves the temporal continuity and trends in the data by estimating missing values based on adjacent time points.
- There are differences in CO, PM10, PM2.5, and NO2 (both in mean and std) indicating a shift in pollutant levels between the two stations.
As per the paper, each experiment is run 5 times with random seeds and the average results are reported. Each iteration has 20 epochs.
The first two tables shown below evaluate the DLinear model. It also shows the results of adding CCM on top of the DLinear model. The input length is fixed at 96.
Aotizhongxin Station
| Horizon | Model | MSE | MAE |
|---|---|---|---|
| 96 | DLinear | 0.653 | 0.520 |
| + CCM | 0.658 | 0.515 | |
| 192 | DLinear | 0.696 | 0.542 |
| + CCM | 0.704 | 0.540 | |
| 288 | DLinear | 0.705 | 0.547 |
| + CCM | 0.704 | 0.544 |
Changping Station
| Horizon | Model | MSE | MAE |
|---|---|---|---|
| 96 | DLinear | 0.729 | 0.538 |
| + CCM | 0.736 | 0.534 | |
| 192 | DLinear | 0.772 | 0.561 |
| + CCM | 0.771 | 0.556 | |
| 288 | DLinear | 0.784 | 0.570 |
| + CCM | 0.780 | 0.562 |
Observations:
CCM slightly increases the MSE compared to the baseline in some cases but consistently lowers MAE. This suggests (1) the dataset likely contains more outliers, and (2) that CCM improves the average prediction accuracy for most channels, but may introduce larger errors in a few outlier cases. However, as shown in the paper, the performance of CCM improves as the forecasting horizon increases.
The better performance on the Changping station may be because Aotizhongxin exhibits more heterogeneous channel behavior which could make clustering less effective.
It may be worthwhile to investigate the impact of explicit outlier removal.
The following table evaluates the zero-shot performance of the DLinear model with and without CCM. Both directions (Aotizhongxin to Changping and Changping to Aotizhongxin) are shown.
| Model Generalization Task | Horizons | DLinear MSE | DLinear MAE | + CCM MSE | + CCM MAE |
|---|---|---|---|---|---|
| Aotizhongxin → Changping | 96 | 0.653 | 0.519 | 0.673 | 0.524 |
| 192 | 0.695 | 0.544 | 0.714 | 0.541 | |
| 288 | 0.704 | 0.547 | 0.720 | 0.560 | |
| Changping → Aotizhongxin | 96 | 0.729 | 0.542 | 0.736 | 0.541 |
| 192 | 0.773 | 0.560 | 0.788 | 0.565 | |
| 288 | 0.782 | 0.568 | 0.805 | 0.576 |
Observations:
The zero-shot results show that DLinear consistently outperforms DLinear+CCM when generalizing from Aotizhongxin to Changping and vice versa.
As highlighted in the dataset comparison, there are significant differences in pollutant distributions across the two stations. These shifts may have disrupted the soft cluster assignments learned during training, causing channels in the target dataset to be matched to suboptimal cluster prototypes.

