The only way to learn math properly is to solve lots of problems. Vereshchagin’s "Collection of Arithmetic Problems: For Secondary Schools' is finally available in English for the first time — 3,238 carefully selected problems with a complete answer key. First published in 1891 by St.
Valeriy M., PhD, MBA, CQF
@predict-addict.bsky.social
Experienced Data Science Leader | PhD in Machine Learning | 7x Author | Black Belt 🥋 in Time Series | Chief Conformal Prediction Promoter| Mathematician | valeriy.ai
**LinkedIn / X (Twitter) combined post:** The only way to learn math properly is to solve lots of problems.
Here is a post optimized for both LinkedIn and Twitter (X). It is structured to be highly scannable, engaging, and clear about your expectations and rewards. ---
Based on the 'Introduction to Conformal Prediction: A Modern, Rigorous Guide with Python' valeman.gumroad.com/... #conformalprediction
Based on the 'Introduction to Conformal Prediction: A Modern, Rigorous Guide with Python' valeman.gumroad.com/... #conformalprediction
Based on the 'Introduction to Conformal Prediction: A Modern, Rigorous Guide with Python' valeman.gumroad.com/... #conformalprediction
A new article 'Hidden Deep Inside a Moscow Museum: A Metal Turtle That Learned' #AI #machinelearning
Last year’s Nation’s Report Card recorded a first: 45% of American 12th graders scored below Basic in mathematics — the highest share ever measured. Only 33% were judged academically prepared for entry-level college math, down from 37% in 2019. The instinct is to blame calculus instruction.
Performing math vs knowing math. That's the difference between modern textbooks and one written in 1884. A.P. Kiselev was a Russian schoolteacher who got tired of watching students struggle with verbose, overcomplicated textbooks that taught procedures without understanding.
Solid mathematical ideas almost always outperform contrived engineering tricks. For years deep learning has been dominated by increasingly complex architectural hacks: CNN blocks, attention layers, channel mixers, residual pathways, normalization stacks.
In forecasting work, I’m continually surprised by how many data scientists interviewing for related roles struggle with something as fundamental as cleaning and filtering data. Interestingly, the great mathematician Andrey Kolmogorov himself created a remarkably effective tool for this…
The Architecture of Weight Decay From Actuaries to DLinear www.youtube.com/watc... Based on the Pro Edition of my book ‘Mastering Modern Time Series Forecasting’ valeman.gumroad.com/...
A Harvard freshman in the new intro calc course passed AP Calculus in high school. She said it felt like “memorising formulas,” not understanding. That’s the difference between AP Calc and Kiselev. Kiselev’s Calculus, first English translation, out now.
Kiselev's Calculus — First Complete English Translation | Russian Math Books
The first complete English translation of Kiselev's calculus — the final volume in his school course. Single-variable differential and integral calculus, proof-based, with classical applications to mechanics and physics. The bridge from Algebra II into the analysis tradition that produced Markov's probability theory.
russianmathbooks.com
One man. Two fields named after him. Dynkin diagrams in algebra. Dynkin's formalism in Markov processes. And in between, E.B. Dynkin wrote a probability book for schoolchildren — Random Walks, straight from the Moscow math circles.
The winner called his solution “Data is Everything.” He was right — twice, and the second time hurts. His extra 600k training rows contained the test labels. His own postmortem admits the contamination. So the interesting question became: how far does his data edge go without the leak?
The hardcover edition of Kiselev’s Arithmetic is gaining on the paperback—and for good reason. Its sturdy, classic cover is built to last, making it a great choice for children.
Still chasing state-of-the-art? Check what your metric actually pays for first. We tested the newest efficiency-optimal conformal methods — CTI (AAAI 2025) and CIR (AAAI 2026), which build provably shortest prediction intervals — against Kaggle’s interval competitions.
The most quoted sentence in probability theory was written by two Soviet mathematicians: "All epistemologic value of the theory of probability is based on this: that large-scale random phenomena in their collective action create strict, nonrandom regularity." Gnedenko and Kolmogorov, 1954.
Based on the 'Introduction to Conformal Prediction: A Modern, Rigorous Guide with Python' valeman.gumroad.com/... #conformalprediction
Hidden deep inside CatBoost is one of the best uncertainty tools almost nobody uses. loss_function="RMSEWithUncertainty". One line. The model fits a Gaussian likelihood — mean and variance jointly — so every prediction carries its own σ. For free. Why should you care?
When it comes to teaching uncertainty quantification, there’s one test most books never take: entering the arena. My new book did. We took its exact workflow — split conformal, CQR, normalised conformal, honest evaluation — into both of Kaggle’s prediction-interval competitions.
Based on the 'Introduction to Conformal Prediction: A Modern, Rigorous Guide with Python' valeman.gumroad.com/... #conformalprediction
Turns out the conformal guarantee costs nothing. We measured it. Twice. Two Kaggle prediction-interval competitions. Birth weight: 690 teams. House prices: 691 teams. We entered both offline with one rule — the exact discipline from my book. Calibration labels touched once.
Based on the 'Introduction to Conformal Prediction: A Modern, Rigorous Guide with Python' valeman.gumroad.com/... #conformalprediction
The winning solution of Kaggle’s house-price interval competition trained on the test labels. Not a rumour. The winner’s own write-up. He found a public King County sales archive with 600k extra rows and trained on it.
Stop using Platt scaling by default. It’s the bug, not the fix. If you’re serious about shipping calibrated models, that one-line reflex is quietly making your best classifier worse. We ran Platt scaling across 21 classifiers, 30 binary datasets, 150 cross-validation folds (arXiv 2601.19944).
If your child learns math like this: 9 × 7 = 63 “just remember it” they will forget it. If they learn why numbers behave the way they do, they never forget. That was the philosophy of Kiselev’s Arithmetic. For 70+ years it defined arithmetic education in Russia.
Dynamic graphs. Temporal fusion. Quantile regression. Conformal calibration. All bolted into one wind-power forecaster. New paper — DG-TFT-CQR — is a kitchen sink of modern forecasting ideas, and it is worth a look for one reason. It ends with conformal calibration. The problem is real.
It's 2026… and for time series forecasting, Transformers & LLMs are STILL what you don't need. This repo has been delivering the truth for couple of years now.