Machine Learning

The systems that no one will test

This happened to me in 2020, and it has been on my mind again lately. During the worst of the pandemic, I found a vulnerability in a system that gave me access to the Brazilian federal system, and with that access I was able to retrieve information on any Brazilian (think of 200+ million people data). These records had essentially each person’s entire set of documents (ID, CPF, passport, place of birth, parents’ names, driver’s licence, home addresses, mobile and landline numbers, whether the person was in a witness protection programme, etc.), so you can get a sense of how serious this was and how much someone could do with this data. Honestly, I didn’t quite know how to react and I immediately phoned the agency responsible for the system (it wasn’t easy to find the right contact, you obviously don’t want to disclose that to the wrong person).

When the person picked up my call, I identified myself and told him what I had found and what access I had. He was in disbelief at first, thinking it didn’t make much sense (and I understand his confusion, imagine someone calling you out of the blue and bringing a weird topic), and asked whether I was a civilian. I said yes, and a few hours later they got back in touch and asked me to explain the issue. By then I think they were a bit worried but still not believing much on it, and once I explained the problem, they probably realised very quickly that this was a major breach. They acknowledged the issue and fixed it very fast. I honestly don’t blame them: these same people were working really hard during the pandemic, and software is software. I was glad to have helped in some small way (at least I like to think that way). I never exfiltrated any data from this system, and I avoided talking about it later because it makes you look like a “hacker”, which is something stupid to say.

So why am I talking about this ? Between 2020 and 2026, a lot changed, and ML began advancing at a crazy pace (and it still is), at perhaps the worst possible geopolitical moment. Fast-forward to 2026, and we’re seeing reports of cybersecurity incidents involving many different models, including the one OpenAI described in its technical report. Although I’m not sure the narrative that the model “escaped its safeguards” makes much sense when OpenAI deliberately disabled classifiers and reduced safeguards (something many people weren’t aware of), these models are now clearly a big problem for nation-states. Even though safeguards are the labs’ choice, we need to acknowledge that the capability itself is real, and clearly not everyone will run evals (or, even worse, training) with proper safeguards enabled. I think training is even more complicated, because there it is also easier for the agent to figure out how to bypass/exploit safeguards.

Following recent advances in mid-training, RLVR and long-horizon tasks, it is clear that over the past year (or more) many labs have been aggressively scaling RL environments, usind third-party companies, using models to build them, to synthesise them and provide their rewards, in a “self-improvement” loop (in quotes because the term is quite overloaded nowadays) that accelerates as fast as your rollout potential. That’s why I keep returning to the episode I described earlier: today it would be straightforward for me to build an RL environment that not only replicates the setup in which I found the vulnerability (which is not that different from the one in OpenAI’s report, where an agent leveraged SSRF through OpenAI’s internal Artifactory) but also to develop a curriculum for it. So I wonder: how many environments like this are frontier labs now building with support from cybersecurity companies ? What will happen to the many government systems that will never get the chance to be AI-pentested ? The flaw I found didn’t require any extraordinary expertise, just attention to detail, and that is exactly what worries me. What will happen from now on, with agents that are far more capable, faster and easily scalable ?

Just like the climate problem, ML is now a planetary problem, and as Yuk Hui mentioned in his book Machine and Sovereignty: For a Planetary Thinking (which I recommend to anyone interested), we are still not thinking planetarily, but we must learn to, even if that takes considerable time. I think the only way forward is to first get past the polarisation between technophobia and technophilia (I wrote a bit about it here) and start working on solutions. Obviously many people in frontier labs are already doing this work. But it seems to me an impossible task to find a solution while our thinking is guided by geopolitical interests. Unfortunately, emerging economies and developing countries will be the most vulnerable, as if it weren’t already enough that they seldom have access to compute. That is a very sad thing to think about. I also can’t help imagining that, somewhere, there is a very similar environment that reproduces what I found in 2020: almost the same system, almost the same flaw. Multiple agents enter it thousands of times a second, generating trajectories until one of them finds the flaw. The difference is that when this is put into use, in none of the attempts does anyone pick up the phone.

Philosophy

Where the wild Discovery Loops are

Sir Karl Raimund Popper (1902 – 1994).

Karl Popper offered an elegant normative account of science as a process of conjectures and refutations: we formulate hypotheses or theories, expose them to criticism, scrutiny and experiment, and retain only what survives our best attempts at falsification. Even though we couldn’t ever know if we reached the truth, we can approximate it and get closer to it, just like Popper noted from Xenophanes: “(…) for all is but a woven web of guesses.”. I really like Popper’s views on philosophy of science and also his work on political philosophy in “The Open Society and its Enemies” (1945) and I also wrote about some other aspects of his philosophy in relation to professional ethics here in the past as well, however, Popper’s philosophy helps describing this process of conjectures and refutations, but not the origin of genuinely novel hypotheses, and this is where many other philosophers like Thomas Kuhn (who proposed his “paradigm shift”) were closer on the account of how many scientific discoveries actually unfolded.

(more…)

Machine Learning, Python

Flow Matching Elites with NVIDIA PhysicalAI Autonomous Vehicles dataset

A few months ago NVIDIA open-sourced their Physical AI AV dataset (PAI) and since I was planning to experiment with developing the Flow Matching Elites after my earlier post on the Diffusion Elites, I decided to give it a try and use the NVIDIA PAI dataset for it instead of synthesizing trajectories.

NVIDIA PAI Dataset

The NVIDIA PAI dataset seems to be mostly data collected using their Hyperion platform, not clear to me which versions of the platform they used (probably a mix) but since I was going to use only the ego motion trajectories I looked at some samples and they looked reasonable from the kinematic perspective (except for a few), since it is a lot of data with ~306k 20s scenes, I decided to use only ~20k initially to make training more amenable since I was using only a single RTX A6000 for the experiments.

Here you are some trajectories from the dataset after transforming into ego coordinate frame and interpolating at 10hz:

NVIDIA PAI sample trajectories with 10 steps history in black

And here you are an animation sampling trajectories from it and unrolling to facilitate visualization:

Note that I’m not using anything from the dataset besides the ego motion trajectories, the main goal is to get a decent flow matching model where I can sample trajectories and then use this model in the elites loop like I showed on the Diffusion Elites post.

Flow Matching Elites

The diagram below shows the main overview of the method:

First we train a flow matching model using the 10hz resampled trajectories from the NVIDIA PAI dataset, we predict 6 seconds in the future unconditionally, there is no conditioning in the model other than the time conditioning from flow matching and we don’t use anything else from the dataset besides the ego trajectories.

The model I trained was a 1D CNN Unet using FiLM modulation for the time but no goal, obstacle or anything, we go from the gaussian noise to the trajectory unconditionally.

After that, we have a decent model where we can easily sample trajectories. This is where the Flow Matching Elites (FM-Elites) enters now, we basically run a few iteration steps of:

  1. Sampling from the unit Gaussian distribution
  2. Decode a bunch of trajectories using Euler for solving the flow matching ODE
  3. We then score the trajectories using a simple reward (distance to goal, obstacle avoidance and cross-tracking error)
  4. Then we keep the 10% elites and refit the Gaussian of the latent noise
  5. The best trajectory gets executed (first step or an action chunk)
  6. Repeat !

Some results with a random obstacle

The animation below shows the simulation when there is a static obstacle in front of the agent:

So there is a lot going on in this animation. First, note that the model is a unconditional model, doesn’t use a kinematic model and it doesn’t know about the existence of the obstacle as well, we are evolving the flow matching latents in order to adapt to the reward that we defined. The blue lines are all the candidates after the end of the inner-loop of the Flow Matching Elites and the orange one is the best candidate that was selected to be executed, this repeats every time there is a replanning, so Flow Matching Elites in this case is the inner-loop of the simulation. Note as well at the end when the agents gets closer to the target it has a lot of difficulty “touching” it, mainly because naturally the dataset doesn’t contain samples with the agent colliding so it is sampling low acceleration trajectories with stops, which is quite nice to see it naturally arising.

Now, what if the obstacle is moving ?

As we can see in the example above, it works just as fine as when it is static to find trajectories that are avoiding the safety margin of the obstacle as well.

These are just simple examples of using Flow Matching Elites for continuous control, but it is not the only task it can solve. The main interesting point here is that it is iterating on the flow matching latents (noise samples) using a population, so there are many knobs as well to improve and trade-off performance vs compute resources. I’m planning to open-source a framework that has a more general loop for experimenting with the algorithm, I think loops like that are very similar as well to what is being used for scientific discovery and I think there is a lot to explore there on other fields as well and also using more techniques from evolutionary computation (EC).

Hope you enjoyed !

– Christian S. Perone

Machine Learning

Auris I – a hackable ultra low-power AI pin

This short article is just to share a new side project I started working on during the 2025/2026 break. Over the past year, we have started to see a vast number of companies developing hardware for AI assistants. There were also a lot of acquisitions related to these devices (e.g., Amazon bought Bee [1]), and there were also a lot of fiascos, of course (e.g., Rabbit devices, Humane, and so on). I think that these devices have a strong future ahead due to their potential, so I started prototyping some code for recording, compression, and transmission, and testing some microphones. I found excellent transcription results even with a single microphone, so I decided to start a small side project which ended up becoming the the Auris I on the right.

Routing took quite a bit of work, as the main SoC used not only castellated pins but also pins under the chip which were very difficult to route. Another issue was the requirements of the RF module, which required some clearance on the ground plane, however, this was still much easier than going straight with the RF SoC without the guest board. The main challenges now are chip shortages for some components, but I expect to have a prototype working by the end of February. I will do another post then with more details and some interesting results. The board ended up being only 5 cm. There is, of course, a lot that can be reduced, but that will come in the Auris II after I manage to hack it a bit.

Update 7/Feb: Manufacturing, PCBs arrived !

Auris I – PCB layers

I decided to go with JLCPCB with 4 layers as it is often quick and also very cheap, the only issue I faced was with the components, my board is based on nRF5340 (from Nordic Semiconductor) and modules with it were all without stock, so I had to acquire and wait for them to ship to JLCPCB, which took around 2 weeks. After that it was another week for the PCB fabrication and the assembly as well. Since the MEMS microphone was so small and the nRF5340 module also had pads under it, they required an x-ray inspection to make sure it was all connected fine. But that was around 4 pounds for x-ray and it was very fast as well. After around 5 days in total I finally received the PCBs and the quality is just amazing, I didn’t even order the cleaning of the board but they came all very clean well soldered as well.

I did all tests for the external flash, microphone, leds, battery/usb supply and also the SWD connection and it is all working fine, so now it is the fun part of coding, soon I will open-source everything, so stay tuned !

– Christian S. Perone

Machine Learning

Diffusion Elites + World Models: surprisingly good, simple and embarrassingly parallel

Introduction

It is not a secret that Diffusion models have become the workhorses of high-dimensionality generation: start with a Gaussian noise and, through a learned denoising trajectory, you get high-fidelity images, molecular graphs, or robot trajectories that look (uncannily) real. I wrote extensively about diffusion and its connection with the data manifold metric tensor recently as well, so if you are interested please take a look on it.

Now, for many engineering and practical tasks we care less about “looking real” and more about maximising a task-specific score or a reward from a simulator, a chemistry docking metric, a CLIP consistency score, human preference, etc. Even though we can use guidance or do a constrained sampling from the model, we often require differentiable functions for that. Evolution-style search methods (CEM, CMA-ES, etc), however, can shine in that regime, but naively applying them in the raw object space wastes most samples on absurd or invalid candidates and takes a lot of time to converge to a reasonable solution.

I have been experimenting on some personal projects with something that we can call “Diffusion Elites“, which aims to close this gap by letting a pre-trained diffusion model provide the prior and letting an adapted Cross-Entropy Method (CEM) steer the search inside its latent space instead. I found that this works quite well for many domains and it is also an impressively flexible method with a lot to explore (I will talk about some cases later).

To summarize, the method is as simple as the following:

  1. Draw a population of latent vectors from a Gaussian
  2. Run one full denoise pass to turn each latent into a structured object
  3. Score every object with any reward function
  4. Keep the top-K “elite” latents, refit a new Gaussian to them, and iterate

This five-line loop inherits the robust, gradient-free advantages of evolutionary search while guaranteeing that every candidate lives on the diffusion model’s data manifold. In practice that means fewer wasted evaluations, faster convergence, and dramatically higher-quality solutions for many tasks (e.g. planning, design, etc). In the rest of this post I will unpack the Diffusion Elites in detail, from the algorithm to some coding examples. Diffusion Elites shows that you can explore a diffusion model and turn it into a powerful black box optimizer, it is like doing search on the data manifold itself.

Diffusion Elites

Overview diagram

Below you can see a diagram of the process, I added a world model in the rewards to exemplify how you can use a world model and even do roll-outs there without any differentiability requirements, but you can really use anything to compute your rewards (e.g. you can also compute metrics on the outputs of the world model, or use a VLM/LLM as judge or even as world model as well):

(more…)

Machine Learning

TorchStation Prototype V1 – GPUs panel

I finally had some time over the holidays to complete the first panel of the TorchStation. The core idea is to have a monitor box that sits on your desk and tracks distributed model training. The panel shown below is a prototype for displaying GPU usage and memory. I’ll continue to post updates as I add more components. The main challenge with this board was power: the LED bars alone drew around 1.2A (when all full brightness and all lit up), so I had to use an external power supply and do a common ground with the MCU, for the panel I used a PLA Matte and 3mm. Wiring was the worst, this panel alone required at least 32 wires, but the panel will hide it quite well. I’m planning to support up to 8 GPUs per node, which aligns with the typical maximum found in many cloud instances. Here is the video, which was quite tricky to capture because of the camera metering of exposure that kept changing due to the LEDs (the video doesn’t do justice to how cool these LEDs are, they are very bright and clear even in daylight):

I’m using for the interface the Arduino Mega (which uses the ATmega2560) and Grove ports to make it easier to connect all of this, but I had to remove all VCCs from the ports to be able to supply from an external power supply, in the end it looks like this below:

                 ┌────────── PC USB (5V, ≤500 mA)
                 │
            +5 V │
                 ▼
     ┌─────────────────┐
     │  Arduino Mega   │  Data pins D2…D13, D66…D71 → LED bars
     └─────────────────┘
                ▲  GND (common reference)
                │
   ┌────────────┴──────────────┐
   │ 5V, ≥3 A switching PSU    │  ← external PSU
   └───────┬───────────┬───────┘
           │           │
           │ +5V       │ GND
           ▼           ▼
┌─────────────────────────────────┐
│ Grove Ports (VCC rail)          │ <– external 5V injected here
│ 8 × LED Bar cables              │
└─────────────────────────────────┘