of all the benchmarks being sent around for Claude 3.7 this is the one i'm paying the most attention to. they're cheating a little bit by giving it the oldest, original pokemon game (red/blue) which is more than 20 years old and will have plenty of info online to learn from.
sheesh! what a day for AI. QwQ-Max, sonnet-3.7 AND open source of FlashMLA
the only interview question i need to ask is: how many sqlite databases do you have on your machine and what do you use them for? if you're able to answer with a specific number, you're not ready yet unless you have a custom cron job cleaning up unused sqlite files.
i don't check social media for 8h and people are already playing with claude 3.7 sonnet
edtech companies who've amassed a bank of their own content in the last decade are sitting on a goldmine.
saving all of my generated deepseek r1's reasoning traces to a USB drive so when the time comes i can send it back to my 12 year old self to prevent him from bashing his head against the desk for being stuck on the 2nd problem of an AMC 12.
DeepSeek, a LLM trained for a fraction of the cost of GPT-Xx models, in 2 months for 6 million, on limited GPUs due to export restrictions, and competing head to head. This is crazy. It's not the AI part I'm excited about, it's the level of efficiency. github.com/deepseek-ai/...
GitHub - deepseek-ai/DeepSeek-V3
Contribute to deepseek-ai/DeepSeek-V3 development by creating an account on GitHub.
github.com
engineering blogs and white papers from the heyday of pre-AI tech are a treasure trove of insights and decision making processes without the burden of validating AI slop. don't discount them merely on their being of outdated.
if you're a cs student aspiring to become a software engineer in 2025, make sure to hone adjacent skills like writing, marketing, & sales. in the age of agents, establishing human connection & building an audience organically will make set you apart from the rest.
supercharge your LLM apps with smolagents 🔥 however cool your LLM is, without being agentic it can only go so far enter smolagents: a new agent library by @hf.co to make the LLM write code, do analysis and automate boring stuff! huggingface.co/blog/smolage...
the intersection between hardcore biotech, precision therapeutics, & AI is an interesting one. i'm particularly bullish about the work being done with organoid intelligence - mainly because of more energy efficient computing as well as possible contributions to neurodegenerative disease research
every PR is another step towards the dream of having a universal near zero latency jarvis+second brain ai mesh seemed like a gargantuan task back when i discovered obsidian/roam in college but current tech & capabilities make it easier to bootstrap a decent representation of this.
the smol.ai newsletter is truly a godsend (thanks @swyx.io) with all of the model releases this week, neurips, and discord chats popping off, having a single place to start from really helps highly recommend.
smol.ai
News and Hackathons for AI Engineers!
smol.ai
when you've been on the internet as long as i have, you would understand that everything can be used for anything and leaves a trail. the difference today is a lot more security theatre & forced transparency it's why i've always written stuff envisioning that they'd be immortalized by a super AGI
annoy an ML engineer with these simple phrases: "cosine distance" "L2 similarity" "but did you ship it?"
My deep learning course at the University of Geneva is available on-line. 1000+ slides, ~20h of screen-casts. Full of examples in PyTorch. fleuret.org/dlc/ And my "Little Book of Deep Learning" is available as a phone-formatted pdf (nearing 700k downloads!) fleuret.org/lbdl/
wanna become a prolific open source contributor? just try using open source llm agent frameworks with rough asynchronous programming patterns in prod
@georgehotz.bsky.social is here! Bluesky is going to be so much fun.
deploying llm apps on modal labs at night is a much better experience than my daytime woes with terraform, cloud build, k8s, and docker
the good thing about this new generation of voice assistants is that everyone can have their own hunter s thompson like narrators for the most mundane parts of their daily routines
i finally asked chatgpt to generate an image given what it knew about me. spot on
Alibaba has their own version on GPT-o1. This might be the best description of “o1-type”systems so far arxiv.org/abs/2411.14405
Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
Currently OpenAI o1 has sparked a surge of interest in the study of large reasoning models (LRM). Building on this momentum, Marco-o1 not only focuses on disciplines with standard answers, such as mat...
arxiv.org
In case you passed out and woke up on saturday lunch. Small models and high quality data are back! ... if they ever left 🤔 - SmolTalk dataset from @huggingface.bsky.social - Tulu 3 models and datasets from @ai2.bsky.social - Nvidia Nymba model from @nvidiastudio.bsky.social
Hello, Bluesky users! I curate and maintain list of resources on testing distributed systems. You might have seen it before. It's a good one, if I may say so myself. asatarin.github.io/testing-dist...
Testing Distributed Systems
Curated list of resources on testing distributed systems
asatarin.github.io
the goodfellow deep learning book is almost a decade old 🤯 2nd edition when?