performance
6 posts found
Fullstack Rust Beyond Electron: An Architectural Case Study of Dioxus
Shipping a 150MB Chromium shell for a basic desktop utility is an outdated architectural pattern. Here is how Dioxus uses Rust signals, system webviews, and Axu…
Model Optimization: From FP16 to Speculative Decoding
A timeline tracking how deep learning model optimization evolved from FP16 mixed precision and INT8 calibration to outlier-aware AWQ, speculative decoding, and …
Your Model Isn’t the Bottleneck. Your Queue Is.
A classifier consumed 0.3% of request time; the queue ate the rest. Why AI deployments fail around the queue, not the model.
Cache Invalidation via Surrogate Keys: Edge Architecture and Purge Mechanics
Decouple cache invalidation from URL paths using Surrogate-Key headers for targeted CDN purging.
Why MoE Models Stream From NVMe: Kernel Prefetch, Read-Ahead, and Async I/O
Frontier MoE models run on consumer hardware by streaming expert weights from disk. Not compression—kernel I/O prefetching overlapped with GPU compute. Here's h…
Kernel Socket Tuning for High-Concurrency Load Balancers: TCP_NODELAY, TCP_FASTOPEN, and Listen Backlog
Disable Nagle, enable TCP Fast Open, tune listen backlog. Three socket parameters cut tail latencies 40–60% on modern load balancers.