What are your favorite recent papers on using LMs for annotation (especially in a loop with human annotators), synthetic data for task-specific prediction, active learning, and similar? Looking for practical methods for settings where human annotations are costly. A few examples in thread ↴
🐘
@pkydrm.bsky.social
research scientist @MosaicML x @Databricks re: rlhf, humans in the loop, and figuring out what it means to have a good model 🤖🧑🎨✨
I am once again pitching my romantic comedy: - two academics start dating - discover they are each other's terrible reviewer - hijinks ensue Working title: Love is Double-Blind
i wish i could shout this from the rooftops. relatedly, there's no need for robots to be limited by the human form. similar/tangential thing came up in the 2010s with respect to self-driving: just because people only sense using their eyes doesn't mean cars have to only use cameras!
A sensible perspective on humanoids in manufacturing (TLDR: if you can make humanoids, you can probably make better, more manufacturing specific things) blog.spec.tech/p/humanoid-r...
we are living in an empirical world and we are empirical girls
A more technical white paper is coming but I learned lots during this process not least of which is that the vast vast majority of RLXF papers over the last couple years are useless. Many assumptions made esp at small scales are simply wrong at larger scales
No labels, no problem! I am so excited for this release. We have been working on it for many months, and it's motivated by a common customer roadblock: insufficient labeled examples.
The hardest part about finetuning is that people don't have labeled data. Today, @databricks.bsky.social introduced TAO, a new finetuning method that only needs inputs, no labels necessary. Best of all, it actually beats supervised finetuning on labeled data. www.databricks.com/blog/tao-usi...
has anyone successfully gotten very involved with their local library system and, if so, how does one do so? i know there are volunteer opportunities and it is my dream to one day organize a crafting circle, but i'm talking about how the library actually organizes / functions / prioritizes things!
🧵 Super proud to finally share this work I led last quarter - the @databricks.bsky.social Domain Intelligence Benchmark Suite (DIBS)! TL;DR: Academic benchmarks ≠ real performance and domain intelligence > general capabilities for enterprise tasks. 1/3
very demure, very mindful, very 2019-era mujoco humanoid learning to walk
here's a Sora generated video of gymnastics
"technology built to address people's needs" is the north star. side note: it would be amazing to see this attitude in the physical, embodied world as well. it's amazing to see how older adults in dense, walkable areas have such different lifestyles than those in car-centric suburbs.
Would love a focus on systems that help older people!!
this is incredible research, and beautiful. would love to know more about what it's like to meaningfully interact with genie 2, or similar models, e.g. to modify the outputs of such a model in the service of a design vision.
Genie 2 can also turbocharge environment design for humans, making it possible to step in and play from concept art 🎨, such as the beautiful work below from one of our rockstar designers.
i often talk about the importance of aligning both the magnitude AND direction of a workstream vector. 1/5
I wrote some thoughts on how to build good LM benchmarks: ofir.io/How-to-Build...
i do not study this, but i did just finish reading the anxious generation and so i'm very grateful that there are so many people who do indeed study such important things!
A start on who to follow for the science on social media and adolescent mental health! Who else is here? go.bsky.app/2PqckAy
When you fail to parse your data that’s a jsonl