Posts

Showing posts with the label MLX

Lunar CLM: From 869 ms to 20 ms per Decision on a Mac

Image
In my last lander post I spent a few paragraphs on Jev , the System One decision model TypeSafe AI opened to early access this month. You hand it a block of state and typed questions, and it returns a probability for every choice. No text to parse. I also made a claim I had not tested: my lander talks to its model through one call, agent.predict(prompt, QUESTIONS) , so swapping in a different decision model should be an adapter, not a redesign. Then I read about Contrastive Language Models (CLM). It takes the same contract and builds it in the open: a block of state in, typed questions, a probability per choice out, with source code and released 8B heads you can run yourself. Its server even answers on an endpoint called /v1/systemone . Jev I can only read about. CLM I could actually test. I had to find out. The adapter part turned out to be true. It was everything after that which kept me busy. The small CLM I trained on my Mac, with no overrides. The lander starts upside do...

The Trainee Got Better. It Still Cannot Fly Alone.

Image
I ended The Engineer and the Trainee with a plan: roll out the shielded pilot, relabel every visited state with MPC, retrain, and see whether the trainee could finally fly on its own. That is still the plan. But before getting there I did something lazier, and the result was more interesting than I expected. I also fixed something that had been bothering me every time I watched a flight. The same recorded flight as last time, re-run with a smoother controller and a retrained model. The engine loses 60% of its thrust at ten seconds, the estimate falls to 40%, and the shield now steps in on about one command in eight instead of one in three. The lander had the shakes Watching a descent, the ship twitched. Not a rendering artifact — the controller really was changing its mind. I measured it: over 30 flights the beam search picked a different command on 77.7% of decisions. Every fifth of a second it re-derived the whole plan from scratch, and when two commands were nearly tie...

The Engineer and the Trainee: Adaptive MPC Guiding and Shielding Laya

Image
Two of my recent lander experiments kept nagging at me. In One Lander, Three Controllers , an adaptive MPC learned its engine during the flight and landed through a power loss with no training data at all. In Lunar Laya , a small decision model on Apple MLX turned a text description of the flight into structured choices, but the text carried hints from a fixed feedback law that assumed a healthy engine. Cut the engine power and that model crashed every time, faithfully following bad advice. So I asked the obvious question: what happens if the engineer writes the hints for the trainee? The result is Lunar MPC + Laya . An actual recorded flight. Laya proposes every command from telemetry alone; adaptive MPC vetoes about one in three. The engine loses 60% of its thrust at ten seconds, the estimate drops to 40%, and the lander touches down at seventy. The engineer and the trainee The game is the same arcade-style lander from Lunar Laya: three pads on bumpy terrain, a tilt control, a...

Lunar Laya: From State and Text to Decisions on a Mac

Image
I keep coming back to Lunar Lander. It gives me a very visible way to ask a simple question: what does a model actually decide? Turn too much, burn too late, or miss the pad, and the result is right there on the screen. After exploring learning and control approaches to this problem, I wanted to try another angle: give a small model the current state and a text instruction, then ask it for a decision that the application can execute and measure. That experiment is now public as Lunar Laya , an Atari-inspired lander using Laya with local inference on Apple MLX. I believe the input should be a state and a text, even a structured text. The output should always be structured and very deterministic, so it can be controlled and measured. This implementation is a representation of that idea. The fact that the model can be small enough to run locally in an MLX environment is proof that it can be done for this kind of bounded decision task. An actual trained-model flight using FP16 MLX ...

FELN Liquid: Fine-Tuning a GIS Model on a Mac

I have been experimenting with small language models that translate everyday questions into structured GIS queries. With FELN Liquid , I wanted to explore the training side on the Mac itself. Could I take a compact model, teach it the conventions of a North Sea catalog, and complete the experiment locally? The first complete run used LiquidAI's LFM2.5-1.2B-Instruct with LoRA through MLX on an Apple M4 Max. Training took 48 minutes and 37 seconds. Including checkpoint evaluation and the final test, the workflow finished in just under an hour. That is a practical turnaround for trying an idea and inspecting what it learned. The task is familiar from the other FELN projects: turn a request into a JSON plan containing layers , where , and relations . A compact Layers.json catalog accompanies each request, supplying the names, types, aliases, codes, and hints that make the data meaningful. For example, “Show all wells with water depth > 350 meters” should produce: { ...