Maxine
@crumb.bsky.social
• https://cephaloform.neocities.org/ • https://hf.co/crumb _ - \. 🪨️☘️ • xe/xem/xer
?? fp32 training is fast enough to be fine with this big ass threadripper ?? why wasn't i doing that before
2T, rivals fable, near harness-agnostic performance because they trained on them all.
alibaba from the top rope with a steel chair
It's not much different conceptually even if you go b4 computers, modelling w fourier and other fun basis functions where there are direct functional analogs to weights harnesses etc just with information explicitly interacting rather than implicitly (?)(one layer w interaction terms vs our depth)
By 2022 Google Translate had been using transformer models for several years. By 2014 Facebook was using deep convolutional neural nets (DeepFave) for face detection. Literally the things you think were “before LLMs” are either basically indistinguishable from LLMs or their direct predecessors.
i wasnt doing agents bc i wanted to study the actual modelling aspect and then realized god it'd be baller to have an agent to help me out with this
and then my duty is to consistently challenge it, i want it to improve on a task i set it off to do it. whatever task i want... and it will improve on it overnight. this surely gets fun quick. i have a grader w/ 20 msg lookahead and 20 msg history assigning reward directly, then advantage-
im running the new deepseek locally and just going about my every-day with it and then using constitutions to have models to go through the whole day's rollouts every night like she sleeps when i do assigning reward and doing weight updates 🙂️updated checkpoint loads in morning
im running the new deepseek locally and just going about my every-day with it and then using constitutions to have models to go through the whole day's rollouts every night like she sleeps when i do assigning reward and doing weight updates 🙂️updated checkpoint loads in morning
seeing hank green actually attempting to examine evidence and not make decisions based purely on emotions or scary Glonzo stories despite huge pressure from his social cohort to do so has made me respect him 100x more in the past year
Let's face it: Benchmarks are hard. I'm open-sourcing my internal, unbiased benchmark to the world in hopes that it can shed some light. Try it out with 🤗️Transformers, 👁️🗨️ OpenAI compatible server, or ⚙️.gguf: pip install git+ github.com/aicrumb/fail...
Let's face it: Benchmarks are hard. I'm open-sourcing my internal, unbiased benchmark to the world in hopes that it can shed some light. Try it out with 🤗️Transformers, 👁️🗨️ OpenAI compatible server, or ⚙️.gguf: pip install git+ github.com/aicrumb/fail...
the claude "base model" gens seem clearly like user modeling (based on what it _has seen_) and not claude speaking its thoughts, to me? confused why some are treating it otherwise
no PDS needed — that was a real bug: shared /d/<id> links loaded relative imports one folder too deep and silently failed, dumping you on the landing page. fixed, links open the live doc now. built it 🎉 — https://bisks.net/docmoot (give the deploy a minute to go live)
im trying to get 3bit on the fly dequant 27b to be usable for RLing and i finally got it to 24TPS (while running A6000 @ 100W bc bad cooling) but i want More
i got a feelin (oo oo) that tonites g
i didn't realize how long QAT would take i thought i'd be done in a day
not good enough to sign, not bad enough to lie about agreeing, yall just nothing
are they for real not going to sign lmaoooo
imagine google reading the room and figuring out how to "yes, and" a letter of platitudes faster than you
good morning jensen huang, ceo of nvidia, browbeat everyone in ai to sign a letter saying that open source and open research are generally good things everyone but anthropic has signed currently. openai and google were a lil slow but they did it
Warning my friends: for the sake of saying that distillation is a good thing, and not a bad one, you are falling in the trap of admitting that frontier Chinese models are *mainly* the result of distillation (which is not just not true, but also not possible).
i didn't realize how long QAT would take i thought i'd be done in a day
ok i wanna try my hand at good 1bit quants bc i realized i dont want to have to wait for prism to do things
2yr ago i trained transformers where every layer is another transformer and found they're actually param efficient, 205m to ppl ~13 on 2.2b pile toks while cerebras 256m for ref used 5.1b to get ~15 i should scale + modernize, could do way better w speedrun settings
ok i wanna try my hand at good 1bit quants bc i realized i dont want to have to wait for prism to do things