Vikram Jha

@invinciblejha.bsky.social

Dedicated to advancing Generative AI, Cybersecurity AI, Advanced Agentic AI Systems, Enterprise SaaS across major underserved industries. Co-Founder @ Stealth AI Startup https://www.linkedin.com/in/invinciblejha/

Rght now American companies need Chinese models to secure their cyber infra due to guardrails on closed models. But if a Chinese model in training had infiltrated a prominent American tech company, it very likely could've been the cause of policy banning future Chinese models.

The goal with my rlhf book is to make the "home on the internet" for the next generation learning post-training. That's why I'm doing all formats (lectures, code, book, discord, model completions... & ofc blog of interconnects). A hub is more lasting than non-fiction writing. Join! rlhfbook.com

Bild

On-policy distillation is on track to be a lasting method in post-training. The list of areas would be: Instruction tuning (SFT/IFT) RLHF Direct Preference Optimization (DPO et al) RLVR On-policy Distillation (OPD) New classes of methods are rare! Excited to play.

Reproducing all of Jürgen Schmidhuber’s papers (1990-2025) using an AI coding assistant. Cool project by Yaroslav! It even reproduced the “World Models” paper by me and Schmidhuber (2018) using a toy environment, with a full VAE + RNN world model implementation. Project: github.com/cybertronai/...

Most obvious immediate impressions on coming back to the US from China. 1. The cars here are so lame. So many cool EVs in China, feels like I went 2 decades back in time. 2. The coffee here is so much better. First real coffee in almost 2 weeks changes you.

🚨Our NeurIPS 2025 competition Mouse vs. AI is LIVE! We combine a visual navigation task + large-scale mouse neural data to test what makes visual RL agents robust and brain-like. Top teams: featured at NeurIPS + co-author our summary paper. Join the challenge! Whitepaper: arxiv.org/abs/2509.14446

Mouse vs. AI: A Neuroethological Benchmark for Visual Robustness and Neural Alignment

Visual robustness under real-world conditions remains a critical bottleneck for modern reinforcement learning agents. In contrast, biological systems such as mice show remarkable resilience to environ...

arxiv.org

NEW: Weekend News Agents No-one can be certain how transformational AI will be:to the economy, to human psychology, to the human experience. I am absolutely certain our politicians haven’t thought about it enough or begun to properly grapple with the economic, regulatory and political implications

Fed up with Meta? Avoiding Instagram or Facebook isn’t enough to stop Meta from harvesting and profiting from your private information. Here’s how to limit Meta’s ability to monetize your personal data.

Mad at Meta? Don't Let Them Collect and Monetize Your Personal Data

If you’re fed up with Meta right now, you’re not alone. Meta tracks you across millions of websites and apps and its business model relies on your data. If you want to limit Meta’s ability to collect ...

eff.org