Bastian Beitzinger

@bstn777.bsky.social

nanomaterials scientists burning tokens. https://github.com/letrplB

Just read about the security incident at openai. I am wondering why Fable has such strict safeguards - it even refused to work on a data schema containing "cell_toxicity", while 5.6 Sol works fine. Both are almost even on capability. And Sam called wolf today for the first time…

Bild

New Mistral model to save EU AI program single-pawedly. Internal codename “le gros chaton“ Napoleon class model with 30T params, 256 experts. Project Glasspaw: Controlled access for critical infrastructure to secure bureaucracy. Makes Claude Mythos look like yesterday’s baguette.

Bild

The opus 4.8 system card reads like a love letter to knowledge work. Honesty and reducing hallucinations seems to be the overarching focus of this release. "Claude Opus 4.8 is the first model to achieve a perfect score on this evaluation-that is, it never reports false numbers."

BildBild

Supply chain attack roulette is back. This time: tanstack. This appears to be the first documented case of a self-propagating ("worm") npm supply-chain attack that produced malicious packages with valid cryptographic signatures and SLSA provenance attestations.

TeamPCP's Mini Shai-Hulud Is Back: A Self-Spreading Supply Chain Attack Compromises TanStack npm Packages - StepSecurity

The Mini Shai-Hulud worm is actively compromising legitimate npm packages by hijacking CI/CD pipelines and stealing developer secrets. StepSecurity's OSS Package Security Feed first detected the attac...

stepsecurity.io

Got a little into Claude Design lately. It makes the design process much more approachable than in a coding environment and has genuinely good taste. The consequences? There are no more excuses to have a bad website anymore. Quality is $20-100 and 4-5 h casual conversation away. Same with slides btw

Milla Jovovich (yes, Leeloo from The Fifth Element) just co-released MemPalace, an open-source, fully local AI memory system that scored the first perfect 100% on LongMemEval (500/500). From kicking zombie ass to building the ultimate agentic memory palace. What a time to be alive.

GitHub - milla-jovovich/mempalace: The highest-scoring AI memory system ever benchmarked. And it's free.

The highest-scoring AI memory system ever benchmarked. And it's free. - milla-jovovich/mempalace

github.com

I found it quite intuitive to let agents detect the limitations and inefficiencies of their own tools and improve based on that. Been successfully using a tailored skill-based workflow for some months now. This tool generalises the concept in the sense of Karpathy's "autoresearch”. Might be useful.

GitHub - kevinrgu/autoagent: autonomous harness engineering

autonomous harness engineering. Contribute to kevinrgu/autoagent development by creating an account on GitHub.

github.com

This is turning into the craziest week in AI infra I can remember. Yesterday Anthropic shipped a .map file in their Claude Code npm package enabling reconstruction of the full source code: ~1,900 files. 512K+ loc. Accidentally open-sourcing one of the best agent harnesses in production.

The fact that every scientific paper in 2026 is still uploaded only as fully formatted PDFs to academic archive sites that often limit downloads tells you everything you need to know about how quickly the scientific system is adjusting to the potential of AI to accelerate science & help discovery.

what a crazy timeline. 5 days ago a new model was leaked and flagged as unprecedented cyber security risk. today one of the operationally most sophisticated supply chain attacks on a top-10 npm package happened. not saying this is connected, but this is what I expect to see more of.

Supply Chain Attack on Axios Pulls Malicious Dependency from...

A supply chain attack on Axios introduced a malicious dependency, plain-crypto-js@4.2.1, published minutes earlier and absent from the project’s GitHu...

socket.dev