Paul Chang

@mummitrollet.bsky.social

ML + stuff @Datacrunch

❗️ We just expanded our capacity of B200 SXM6 180GB servers – available in the DataCrunch Cloud Platform. The best thing is… You can deploy the Blackwell platform without approvals. Just sign in, select the instance type, and start your deployment: cloud.datacrunch.io?utm_source=b...

Bild

🆕 Inference API for FLUX.1 Kontext [max] & [pro] are now available on DataCrunch! We are an infrastructure partner of Black Forest Labs for Kontext, a suite of generative flow matching models for text-to-image and image-to-image editing. Learn more: datacrunch.io/managed-endp...

Bild

🚨 Summer Inference by Symposium AI is happening next Wednesday, June 4, at 16:00-22:00. 🇫🇮 This event will bring together 250 AI engineers, researchers, and founders under one roof in Helsinki. 🔗 You can still grab one of the last remaining seats: lu.ma/x5hhj79x

Symposium AI - Summer Inference · Luma

Join 250 leading AI builders for an epic night in Helsinki! Symposium AI events bring together top AI talent, researchers, and engineers who are actively…

lu.ma

I don’t mean to be a broken record but AI development could stop at the o3/Gemini 2.5 level and we would have a decade of major changes across entire professions & industries (medicine, law, education, coding…) as we figure out how to actually use it & adapt our systems. AI disruption is baked in.

BildBildBildBild

1/10🔥 New paper alert in #AABI2025 Proceedings! Normalizing Flow Regression (NFR) — an offline Bayesian inference method. What if you could get a full posterior using *only* the evaluations you *already* have, maybe from optimization runs?

Bild

1/ We asked GPT-4.5 -- allegedly the model with the best sense of humor, according to the site we do not talk about here -- to write a comic about our recent AISTATS paper on the Amortized Conditioning Engine (ACE). Then gpt-4o drew it. You judge the result... (text continues 👇)

A comic "Bayes explains everything!"

The Llama 4 model that won in LM Arena is different than the released version. I have been comparing the answers from Arena to the released model. They aren't close. The data is worth a look also as it shows how LM Arena results can be manipulated to be more pleasing to humans. t.co/rqAey9SMwh

BildBildBildBild

I wanted to change the color scheme on a blog for some plots so I decided to test Claude code. github.com/datacrunch-r... I didnt relaize it inserts "Co-Authored-By: Claude <noreply@anthropic.com>" I would have got away with it if it wasn't for that pesky Claude code.

Update plot-script.py to use 2025 brand colors · datacrunch-research/blogs@6955de0

- Added brand colors 2025 palette - Updated all plots to use the new color scheme - Regenerated all plot images with the new colors - Set consistent style across all plots 🤖 Generated with [Claude...

github.com

This is a new blog looking at the individual optimizations that went into serving the DeepSeek model class in SGLang. I have been observing the SGLang repo for a few months now, and it's crazy how quickly they integrate new optimized features. It's a very cool open-source project!

@datacrunch.io · last yr.

New blog post: Optimization techniques applied by the SGLang team for DeepSeek-V3 inference. You'll find a comprehensive overview of the techniques, their benefits and implications, and our benchmarks. datacrunch.io/blog/deepsee...

Llama 4 uses both interleaved chunked attention and global (NoPE) attention mechanisms, similar to a recent Cohere paper. It's cool to see innovation in attention layer architectures for the large models, and it showed to the world. arxiv.org/abs/2501.18795.

Rope to Nope and Back Again: A New Hybrid Attention Strategy

Long-context large language models (LLMs) have achieved remarkable advancements, driven by techniques like Rotary Position Embedding (RoPE) (Su et al., 2023) and its extensions (Chen et al., 2023; Liu...

arxiv.org

🚨 NVIDIA HGX B200: available NOW on DataCrunch! Be among the first to gain instant access to 1x, 2x, 4x, and 8x B200 GPUs with our high-performance VMs. Sign up and enjoy expert support with secure service where performance meets sustainability. 🔗 cloud.datacrunch.io

Bild

I just came across this paper from ICLR 2024 which proposes an intricate combination of transformers and diffusion models to generate forecasts with these uncertainty bounds (red box), which are clearly inappropriate and could likely be outperformed by modelling the data as a random walk...

A grid of plots of time series forecasts, the proposed model's forecasts are highlighted by a red box. The uncertainty estimates for the proposed model are constant over the forecast horizon and are poorly calibrated.

🥉 SemiAnalysis awarded DataCrunch with bronze on the GPU Cloud ClusterMAX™ Rating! We thank their team for this independent evaluation, validating our approach to pushing the boundary of resource-efficient AI infrastructure ⬇️

Bild

I've been diving into the "black magic" world of CUDA recently. More posts may follow, but I think we're at an interesting point. Perhaps the "CUDA moat" is under pressure and perhaps changing how we interact with GPU programming. 🧵