A kitchen-sink collection of experimental auxiliary-head architectures built on DeepSeek Flash, exploring anything that might extend MOE architectures
Nicholai Mitchko
nmitchko
AI & ML interests
Fine-tuning, Scaling, Enablement, Activation Engineering
Recent Activity
published a model 2 days ago
nmitchko/DeepSeek-V4-Flash-0731-Latent-Reasoning updated a model 2 days ago
nmitchko/DeepSeek-V4-Flash-0731-Latent-Reasoning upvoted a paper 14 days ago
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning
ChainsOrganizations
None yet