El Peso del Conocimiento: visualicé la arquitectura de Kimi K3 y aquí está para ti
Bajé pesos reales del modelo de IA de pesos abiertos más grande publicado hasta hoy y los conté. Son 21 cifras.
El 27 de julio de 2026 Moonshot AI publicó los pesos completos de Kimi K3: 2,8 billones de parámetros, 1,56 TB, licencia de tipo MIT. Moonshot lo presenta como el primer modelo abierto de la clase de los 3 billones, y no encontré uno mayor: el primero de ese tamaño que cualquiera puede descargar. Casi todo lo que se escribió sobre él repitió las mismas cifras de prensa. Yo quise hacer otra cosa: abrir el archivo y mirar.
Un parámetro es una perilla. Durante el entrenamiento, el modelo giró billones de perillas millones de veces hasta que sus respuestas parecieron inteligentes. Al usarlo, quedan congeladas. Eso es lo que se descarga: un archivo con 2,8 billones de números y nada más. Ni reglas, ni base de datos, ni textos guardados adentro.
Para sentir el tamaño: si leyeras un número por segundo, sin dormir, tardarías 88.727 años. Si los imprimieras a 3.000 por página, la pila mediría 93 kilómetros, diez veces el Everest. Y todo eso cabe en un disco de cien dólares.
Lo que hicimos distinto
Los archivos del modelo tienen una tabla de contenidos que dice en qué bytes vive cada pieza. Se puede pedir al servidor solo esos bytes. Así bajamos un experto completo, 33 millones de números en tres matrices, tocando el 0,1% de un archivo de 17 GB, y después una matriz por cada una de las 92 capas con expertos. Unos 600 MB de 1,56 TB, el 0,04% del modelo. Los decodificamos con una implementación verificada contra la de referencia de gpt-oss y los contamos, uno por uno.
21 cifras
Una matriz típica de un experto tiene 11.010.048 pesos y usa exactamente 21 valores distintos. No 21 millones. Veintiuno. En 55 de las 92 capas son exactamente 21; las primeras capas llegan a 53. Es consecuencia del formato de 4 bits con que se guardan: 8 magnitudes, 2 signos y unas pocas escalas por bloque de 32 números. Medio byte por peso.

Conteo real sobre la matriz w1 del experto 733 de la capa 45. Ámbar empuja a favor, rojo en contra.
Y aquí viene lo que me cambió la cabeza. Comparé expertos de una misma capa y sus estadísticas son casi idénticas: misma proporción de ceros, misma magnitud promedio, diferencias menores al 3%. Pero no son intercambiables: la similitud de coseno entre dos expertos vecinos va de 0,000 a 0,002 en las tres capas donde la medimos. Tienen los mismos ladrillos y apuntan a direcciones sin relación alguna. La inteligencia no está en la riqueza de los números individuales. Está en su arreglo.
Lo que esto no es
Es un escáner del tejido, no una lectura de la mente. Ningún número guarda un hecho: "Santiago es la capital de Chile" no vive en una celda que puedas señalar; vive repartido como patrón entre miles de millones de cifras. Medimos 95 de 82.432 expertos, una muestra y no un censo, y cada número del atlas dice de dónde salió. Tampoco tocamos todo el modelo: K3 además ve, con un codificador de imágenes de 401 millones de parámetros que aquí no aparece. Y estas 21 cifras no son una compresión posterior: el modelo se entrenó sabiendo que viviría en 4 bits. El script y los datos están publicados para que cualquiera pueda comprobarnos.
Para quien decide
Inteligencia de frontera convertida en archivo. Lo que hace dos años solo existía detrás de una API hoy se descarga con una licencia de tipo MIT, permisiva salvo para quien opere modelo-como-servicio a gran escala; la pregunta dejó de ser si tu organización tendrá acceso y pasó a ser qué hará con él. El costo visible baja y los invisibles suben: el peso ya no está en el proveedor sino en tu capacidad de correrlo, ajustarlo y entender sus límites. Para una empresa en Chile o en América Latina eso es, a la vez, una oportunidad de soberanía técnica y una dependencia nueva, de talento y de cómputo en vez de contratos.
Y lo que la medición enseña sobre cómo usarlos: adentro no hay archivos secretos ni verdades guardadas. Hay 21 cifras repetidas billones de veces en arreglos que nadie diseñó. Son instrumentos estadísticos poderosos, no oráculos. Diseñar con esa humildad es la diferencia entre adoptar y depender.
Entra al atlas
Todo esto se puede recorrer. Del modelo completo a un número real en cuatro niveles de zoom, con las mediciones a la vista: la historia, el atlas interactivo, los pesos, de verdad (una matriz completa decodificada en tu navegador desde los bytes del modelo) y qué estoy viendo, cada término explicado en simple.
Pesos: Kimi K3 © 2026 Moonshot AI, bajo la Kimi K3 License. El atlas publica mediciones y cuatro matrices completas, una por experto, de los 82.432 expertos; no redistribuye el modelo.
The Weight of Knowledge: I visualised Kimi K3's architecture and here it is for you
I downloaded real weights from the largest open-weight AI model published to date and counted them. They are 21 values.
On July 27, 2026, Moonshot AI released the full weights of Kimi K3: 2.8 trillion parameters, 1.56 TB, an MIT-style license. Moonshot presents it as the first open model of the 3-trillion class, and I found none larger: the first of that size anyone can download. Almost everything written about it repeated the same press figures. I wanted to do something else: open the file and look.
A parameter is a knob. During training, the model turned trillions of knobs millions of times until its answers looked intelligent. When you use it, they are frozen. That is what you download: a file with 2.8 trillion numbers and nothing else. No rules, no database, no stored texts inside.
To feel the size: reading one number per second without sleeping, you would need 88,727 years. Printed at 3,000 per page, the stack would stand 93 kilometres tall, ten Everests. And all of it fits on a hundred-dollar drive.
What we did differently
The model files carry a table of contents that says which bytes hold each piece. You can ask the server for just those bytes. That is how we pulled one complete expert, 33 million numbers across three matrices, touching 0.1% of a 17 GB file, and then one matrix from each of the 92 layers with experts. About 600 MB of 1.56 TB, 0.04% of the model. We decoded them with an implementation verified against the gpt-oss reference and counted them, one by one.
21 values
A typical expert matrix holds 11,010,048 weights and uses exactly 21 distinct values. Not 21 million. Twenty-one. In 55 of the 92 layers it is exactly 21; the earliest layers reach 53. It is a consequence of the 4-bit format they are stored in: 8 magnitudes, 2 signs and a few scales per block of 32 numbers. Half a byte per weight.

A real count over matrix w1 of expert 733, layer 45. Amber pushes with, red against.
And here is what changed my mind. I compared experts within one layer and their statistics are nearly identical: same share of zeros, same average magnitude, differences under 3%. Yet they are not interchangeable: the cosine similarity between two neighbouring experts runs from 0.000 to 0.002 in the three layers where we measured it. Same bricks, pointing in completely unrelated directions. The intelligence is not in the richness of the individual numbers. It is in their arrangement.
What this is not
It is a scan of the tissue, not a reading of the mind. No number stores a fact: "Santiago is the capital of Chile" does not live in a cell you can point at; it lives spread out as a pattern across billions of numbers. We measured 95 of 82,432 experts, a sample and not a census, and every number in the atlas says where it came from. Nor did we touch the whole model: K3 also sees, through a 401-million-parameter image encoder that does not appear here. And these 21 values are not a later compression: the model was trained knowing it would live in 4 bits. The script and the data are published so anyone can check us.
For those who decide
Frontier intelligence turned into a file. What two years ago only existed behind an API now downloads under an MIT-style license, permissive except for those running model-as-a-service at scale; the question is no longer whether your organization will have access but what it will do with it. The visible cost drops and the invisible ones rise: the weight is no longer on the vendor but on your ability to run it, tune it and understand its limits. For a company in Chile or across Latin America that is, at once, an opportunity for technical sovereignty and a new dependency, on talent and compute instead of contracts.
And what the measurement teaches about using them: inside there are no secret files and no stored truths. There are 21 values repeated trillions of times in arrangements nobody designed. They are powerful statistical instruments, not oracles. Designing with that humility is the difference between adopting and depending.
Enter the atlas
All of it can be explored. From the whole model to one real number in four zoom levels, measurements in plain sight: the story, the interactive atlas, the weights, for real (one complete matrix decoded in your browser from the model's own bytes) and what am I seeing, every term explained simply.
Weights: Kimi K3 © 2026 Moonshot AI, under the Kimi K3 License. The atlas publishes measurements and four complete matrices, one per expert, out of 82,432 experts; it does not redistribute the model.