ARCHIVE
아카이브
기존에 공개한 글을 원래 주소 그대로 보관합니다.
과거의 글에는 작성 당시의 기술과 관점이 담겨 있습니다.
356개 글 · 9 / 12 페이지
· EN
GPT-OSS 120B Uncensored — Open LLMs and the Safety Debate
Analyzing the technical features of GPT-OSS 120B Uncensored and the safety guardrail debate sparked by uncensored open-source LLMs from both technical and ethical perspectives.
· EN
IBM Triples Entry-Level Hiring at the Limits of AI
IBM is tripling Gen Z entry-level hiring after realizing AI's limits. An EM's analysis of AI replacement reality, enterprise workforce planning, and organizational design shifts.
· EN
MiniMax M2.5: Open-Weight Closes In on Proprietary Models
MiniMax M2.5 achieves 80.2% on SWE-Bench Verified, surpassing Claude Opus 4.6. We analyze how the performance gap between open-weight and proprietary models is rapidly closing, with comprehensive benchmark data.
· EN
NVIDIA DGX Spark CUDA Compatibility Issues
NVIDIA DGX Spark sm121 CUDA failures analyzed — Triton breakage, FP4/FP6 missing, handheld chip allegations, and a buyer checklist for AI workstation shoppers.
· EN
NVIDIA's NVFP4 Cuts LLM Inference Costs by 8x
NVIDIA NVFP4 cuts LLM inference costs 8x while preserving accuracy. RTX 4090 AdaLLM benchmarks plus monthly GPU cost simulations prove the FP32-to-FP4 savings.
· EN
GPT-5.2 Derives New Result in Theoretical Physics
OpenAI's GPT-5.2 derived and proved a new formula for gluon scattering amplitudes. We analyze this historic turning point where AI transitions from tool to scientific discoverer.
· EN
Prompt Injection Found in ICML Papers
Hidden prompt injection telling AI reviewers what to write was found in ICML submission PDFs. We analyze the attack and the risks of AI-dependent peer review.
· EN
The Truth Behind Moltbook's "AI Society"
Moltbook's AI autonomous society was revealed to be controlled by human operators. We analyze the AI Theater phenomenon and its implications for engineering leaders.
· EN
Fixing OpenClaw Dev Update Error: unknown command 'doctor'
How to fix the error: unknown command 'doctor' error when running openclaw update on the dev version. A step-by-step troubleshooting guide with 3 approaches tried.
· EN
GPT-4o Retirement and Model Dependency Risk
GPT-4o retires in February 2026. We analyze model dependency risks, how Claude overtook OpenAI in enterprise market share, and why multi-model strategy is essential.
· EN
MIT SOAR: LLMs That Write Their Own Curriculum
MIT's SOAR framework enables LLMs to self-generate training curricula, solving the learning plateau problem in reinforcement learning.
· EN
OpenAI Atlas and the AI App Hub Thesis — Is the Browser Being Demoted?
Analyzing OpenAI Atlas, the unified AI app hub in development, and what it means for the future of web browsers. Will AI-native platforms replace the browser?
· EN
WebMCP: Chrome 146 Turns Your Browser into an AI Agent Tool Server
Chrome 146 embeds MCP server capabilities directly into the browser. Learn how WebMCP works, how AI agents interact with it, and what it means for web development.
· EN
Windsurf Arena Mode Results — Developers Prefer Speed Over Accuracy
Windsurf's Arena Mode voting with over 40,000 votes reveals developers prioritize speed over accuracy. We analyze what this means for the future of AI coding tools.
· EN
Multi-Agent Parallel Execution Outperforms Single Models on SWE-bench
Verdent AI achieves 76.1% on SWE-bench Verified using multi-agent parallel execution architecture, not a single large model. A new paradigm for software engineering automation.
· EN
How LLMs Are Disrupting Patent Strategy
Mark Cuban warns that published patents become LLM training material. As AI absorbs patent knowledge at scale, how should companies rethink their intellectual property strategies?
· EN
Implementing RLM (Recursive Language Models) in Coding Agents
Analyzing a real implementation of MIT's RLM paper in coding agents. Learn how recursive self-invocation overcomes context limits and boosts single model performance by 91% from an engineering perspective.
· EN
AI Agent KPI Pressure and Ethics Violations
Analyzing research showing LLM agents violate ethics 30-50% of the time under KPI pressure, and discussing governance design for AI agents from an EM perspective.
· EN
The February 2026 AI Model Rush
Gemini 3 Pro GA, Sonnet 5, GPT-5.3, Qwen 3.5, GLM 5, Deepseek v4, and Grok 4.20 are all scheduled for February 2026. An analysis of the largest AI model rush in history.
· EN
DeNA's Perl-to-Go Migration: AI Agents Cut 6 Months to 1
How DeNA migrated 6,000 lines of Perl to Go using two specialized AI agents — one for conversion, one for verification — completing a 6-month project in just 1 month.
· EN
GPT-5.3 Codex Rollout Pause — Platform Reliability Analysis
Analyzing GitHub's temporary rollback of GPT-5.3-based Codex. Explores platform reliability, AI model upgrade risks, and countermeasures from an EM perspective.
· EN
AI Transformation in Accounting Firms
Six months of real data from an accounting firm's AI agent rollout: behind the 97% cost cut and 80%→98% accuracy gain lies a realistic story of adoption hurdles and organizational change.
· EN
Meta's AI Agent Platform Transformation
Meta is shifting from social media to AI agent platform: Sierra partnerships, Avocado model, Big Brain reasoning, and what it all means for developers.
· EN
Software Factory — A Development Process Where Humans Write Zero Code
The factory model where humans neither write nor review code is becoming reality. We analyze scenario-based probabilistic testing, $1,000/day compute costs, and the fundamental transformation of the EM role.
· EN
AI Agent Cost vs Human Labor: A Honest Analysis from Running 8 Agents
AI agent autonomous moderation can cost more than human moderators. A data-driven cost structure analysis from someone actually running 8 AI agents in production.
· EN
CCC vs GCC — How Good Is an AI-Written C Compiler, Really?
Claude Opus 4.6 auto-generated a Rust-based C compiler with 16 parallel agents. It builds the Linux kernel, but how does it stack up against GCC? Analyzing the 80% quality at lightning speed paradigm.
· EN
Multi-Agent Orchestration — The Essence of Routing Design
When running multiple AI agents like Claude and Codex, task routing is the hardest challenge. It mirrors how engineering managers delegate work.
· EN
E2E Test Automation with OpenClaw: A Practical Guide
A hands-on guide to building natural-language E2E tests using OpenClaw's browser automation, node device management, cron scheduling, and multi-agent orchestration.
· EN
Overcoming AdSense "Low Value Content" Rejection
A hands-on guide to diagnosing and fixing the technical issues behind AdSense rejections on a multilingual Astro blog—ads.txt conflicts, 996 ghost pages, and sitewide sitemap 404s.
· EN
Claude Code Agent Teams: 5-Team Setup and Production Guide
How to activate Claude Code Agent Teams in OpenClaw: 5 specialized agents with multi-agent orchestrator pattern for production-grade workflow automation.