Blog
Welcome to the HappyRock blog!
Here we share technical insights, project updates, and industry trends.
Latest Articles
Want to contribute an article? Contact us: info@happyrock.cloud
MMProLong Long Document LMM Training Paradigm Deep Dive: Why QA Pairs Outperform OCR Transcription by 100x — ByteDance Seed Team and HKUST Joint Research
Friday, July 31, 2026 in Blog
1. Introduction: The Hidden Cost of Long-Document Multimodal Training In late July 2026, ByteDance’s Seed team and Hong Kong University of Science and Technology released MMProLong — a framework that breaks through the efficiency barrier of …
LongStraw Long Context Training Breakthrough Deep Dive: 2M Token RL Training on 8 H20 GPUs — Fudan MindLab Memory Wall Breaker
Friday, July 31, 2026 in Blog
LongStraw Long Context Training Breakthrough Deep Dive: 2M Token RL Training on 8 H20 GPUs — Fudan MindLab Memory Wall Breaker 1. Introduction: The Most Absurd Gap in AI Training In July 2026, arXiv:2607.14952 quietly appeared — Fudan University and …
EvoLib Test-Time Learning Deep Dive: Microsoft Gradient-Free Knowledge Evolution with IG and Future IG Credit Assignment
Friday, July 31, 2026 in Blog
EvoLib Test-Time Learning Deep Dive: Microsoft Gradient-Free Knowledge Evolution with IG and Future IG Credit Assignment 1. Introduction: The Post-Deployment Learning Dilemma On July 30, 2026, Microsoft Research open-sourced EvoLib—a Test-Time …
Spotter AI Unlearning Blind Spots Deep Dive: Over-Unlearning and Prototypical Relearning Attack — ICML 2026 Machine Unlearning Full Analysis
Thursday, July 30, 2026 in Blog
Introduction Machine Unlearning (MU) aims to make AI models “forget” specific training data without costly retraining. But a July 2026 study accepted to ICML 2026 reveals two blind spots that have been overlooked: Over-unlearning: …
SenseNova-Vision Unified Vision Model Deep Dive: One Model for Detection, Segmentation, Depth, 3D Reconstruction — Zero Architecture Changes, Surpassing All Specialized Models
Thursday, July 30, 2026 in Blog
Introduction Computer vision has long suffered from the “one task, one model” fragmentation — DETR for detection, SAM for segmentation, MoGe for depth, VGGT for 3D reconstruction. Each model has a different architecture, different data …
Reinforced Dreamer Asymmetric World Model Deep Dive: Fixing Privileged Information Representation Failure with Latent Guidance
Thursday, July 30, 2026 in Blog
Introduction World models are the core technology that lets RL agents “simulate the future in their minds.” The Dreamer family of algorithms learns an implicit model of the environment, enabling agents to plan actions in imagination. They …
Moonshot AI Transformer Foundation Rebuild Deep Dive: MoonShadow Optimizer, KDA Linear Attention, and Attention Residuals Reshape LLM Training
Thursday, July 30, 2026 in Blog
Introduction In late July 2026, Moonshot AI founder Yang Zhilin delivered a talk at GTC2026 that reverberated through the AI engineering community. The core message was deceptively simple: to make open-source models match closed-source ones, scaling …
NVIDIA AdamW Optimizer Scaling Ceiling Deep Dive: How Muon/SOAP Become the New Foundation for Trillion-Parameter Training
Wednesday, July 29, 2026 in Blog
1. Introduction: The “Scaling Ceiling” of AI Training Optimizers On July 28, 2026, NVIDIA disclosed research findings that could reshape large model training: when batch size scales to 100 million tokens, the dominant AdamW …
NVIDIA 50B Investment in SSI Deep Dive: Vera Rubin Platform, AI Financing Circularization and Safe Superintelligence Architecture
Wednesday, July 29, 2026 in Blog
1. Introduction: The Technical Logic Behind the Largest AI Financing Deal On July 28, 2026, NVIDIA announced a ~$5 billion investment in Ilya Sutskever’s Safe Superintelligence (SSI)—the largest single equity investment in the current AI boom. …
MCP Protocol 2026-07-28 Stateless Architecture Upgrade Deep Dive: Anthropic Model Context Protocol's Biggest Architectural Revision
Wednesday, July 29, 2026 in Blog
1. Introduction: A Historic Turning Point for MCP On July 28, 2026, Anthropic’s Model Context Protocol (MCP) released the 2026-07-28 specification—the largest architectural revision since the protocol’s inception. The core change …