Greg Faletto

@gregoryfaletto.com

Statistician, Data Scientist.

As many standardized-test-takers know, there are two ways to score highly on a test: learn the material, or guess what the question-writer is thinking. LLMs have gotten really good at the second part!

Grace@gracekind.net · 6mo ago

There’s some cool recent research on this phenomenon! It turns out vision language models excel at image benchmarks *even when the actual images aren’t provided,* because the answers are implicit in the questions! arxiv.org/abs/2603.21687

Can we all agree that naming your ice cream flavor a multisyllable made-up word that includes an intentional misspelling of "mocha" is insufficient notice that it contains a nontrivial amount of caffeine

"We are projecting onto an astral plane of heretofore unknown quality, with perfect genius that God himself could ne'er have foreseen." --an LLM after a couple rounds of iteration to get code that actually runs and does what it's supposed to do

It sort of feels like if you pay $8 million for a 30 second Super Bowl ad, they should throw in the rights to refer to it as the Super Bowl instead of the "big game"

“Promotional image with the text ‘Watch the Alexa+ big game ad’ overlaid on a close-up of a person’s face in a blue-toned scene.”

So it seems like one of the highest-leverage things you can do is learn one of VS Code/Positron/Antigravity/Cursor/Windsurf. Because then you pretty much know the rest of them too

As a person who does believe capable AI risks are real, it’s frustrating to see the community not exactly behaving responsibly. If you can’t find it in yourself to say or do something about a literal, present, AI harm (Grok being used for abuse) what is the point of you

📣We're hosting a rainbowR conference, to connect and promote LGBTQ+ people in the #RStats community and to showcase examples of working with LGBTQ+ data 🌈🎉 To help us plan, if you'd be interested in attending or speaking, please fill in this very short form (1-2 mins) docs.google.com/forms/d/1Tx0...

Interested in a rainbowR Conference?

rainbowR (https://rainbowr.org) is hosting a virtual conference (tentative date: February 26, 2026). The aim of the conference is to connect and promote LGBTQ+ people in the R community, and to showca...

docs.google.com

Starting to think a big divide in whether or not people like AI is: "do you do the kind of work where it's much easier to check and fix an 80% reliable solution than it is to just do it yourself from scratch?"

this rules. Haven't taken the plunge to pay for Gemini access, been making the most of my free daily 2.5 Pro prompts for math. The CLI preview access is a super helpful workaround. I just put my prompt in a text file along with the pdf I had uploaded in a question I had asked in the web interface.

Screenshot of a dark-themed macOS Terminal window running “Gemini.” Across the top, large ASCII-art letters spell “GEMINI” in a pastel gradient from blue (left) through purple to pink (right). Beneath, white text on black lists four “Tips for getting started,” such as asking questions, being specific, creating a GEMINI.md file, and using /help. The prompt then shows a user command instructing Gemini to load a PDF on Oracle efficient variable selection and a question.txt file for feedback. A bordered status box follows, with a green checkmark and the message “ReadManyFiles Will attempt to read and concatenate files using patterns…,” indicating the files are being processed from the user’s Documents directory.A terminal window displaying feedback on a mathematical proof. The text comments on the generalization of Theorem 1 from Kock (2013), noting that the new assumptions are reasonable within cited literature. It includes a detailed breakdown under the heading "Proof Structure and Logic." Step 1 discusses bounding an error term and using assumptions to eliminate a cross-term, and confirms the validity of the argument. Step 2 begins with a comment on a modification of Kock's Lemma 1. The content includes inline LaTeX code for mathematical notation.

Linear regression can be seen as overly simplistic when we have ML models and semiparametric statistical theory to back them up. Linearity often doesn't hold. But linear regression is still useful even in causal inference, where normally we're very cautious about unrealistic assumptions. 🧵

Fully aware I'm touching a hot stove... am I the only one who's not sure that mandating a higher percentage of NIH funds go directly to research instead of overhead is a bad idea? I'm familiar with R1 research environments and how funding works, and I support more funding for basic research.