Lunar CLM: From 869 ms to 20 ms per Decision on a Mac

In my last lander post I spent a few paragraphs on Jev, the System One decision model TypeSafe AI opened to early access this month. You hand it a block of state and typed questions, and it returns a probability for every choice. No text to parse. I also made a claim I had not tested: my lander talks to its model through one call, agent.predict(prompt, QUESTIONS), so swapping in a different decision model should be an adapter, not a redesign.

Then I read about Contrastive Language Models (CLM). It takes the same contract and builds it in the open: a block of state in, typed questions, a probability per choice out, with source code and released 8B heads you can run yourself. Its server even answers on an endpoint called /v1/systemone. Jev I can only read about. CLM I could actually test. I had to find out.

The adapter part turned out to be true. It was everything after that which kept me busy.

A small trained CLM rights an upside-down lunar lander, crosses two hills and lands on the tallest pad

The small CLM I trained on my Mac, with no overrides. The lander starts upside down and at rest over the left hill, rights itself, crosses two hills and lands on the tallest and narrowest pad. That is 639 decisions over 128 simulated seconds, shown at about 8× speed. Every decision matched the guidance request.

Why I keep asking for structured decisions

This is not a new drum for me. Since Lunar Laya I have argued that a model making decisions should take state as input, maybe with some structured text, and return a structured answer: a fixed set of choices, a probability for each, and a deterministic rule for picking one. Jev and CLM arrive at that same contract from two different directions, one as a hosted product and one as open research. Seeing two separate teams converge on it is a big part of why this experiment happened.

The reason I keep pushing is measurement. When every decision is a typed choice made from a known state, I can log it next to what should have happened, count the agreement, time it, replay it and point at the exact frame where a flight went wrong. A free-form answer gives me prose to interpret. A structured answer gives me a number I can compare across runs, models and hardware.

And once a decision can be quantified, it can be improved on purpose. Find the states where it was wrong, relabel them, retrain, and measure again. That loop is how a system gets better by itself, and it only closes because the output is structured. The trainee posts are one turn of that loop, and this post is another. Every number below comes from comparing a typed decision against a reference.

The real thing was slow

CLM works differently from Laya. It embeds the flight state once per question, embeds every candidate action description separately, runs both through trained projection heads and picks the closest match. The action side never changes, so the official server caches it. I installed the official contrastive-lm 0.1.0 engine with its released heads on my Apple M4 Max, had llama.cpp serve the Qwen3-8B embeddings on Metal, and pointed the same lander at it.

The median decision took 869 ms. The lander asks for a decision every 0.2 seconds of simulated time, and CLM met that deadline on none of the 96 states I sampled. It also agreed with both guidance requests on only 15 of them. On one complete flight, raw CLM left the engine off and hit the pad too fast after 93 decisions. Those 18 simulated seconds took 74 seconds of wall time.

Is that fair to CLM? Not entirely. I ran a Q8 quantized encoder through llama.cpp, not the official BF16 setup on vLLM and a GPU. The model had also never seen this task, while the Laya checkpoint I compare against was trained for exactly this job. But the time is not a mystery. Every decision pushes two question-conditioned states, 250–266 tokens in total, through an 8B encoder. Skipping text generation does not skip the encoder.

Keep the idea, shrink everything else

What I like about CLM is the shape of it: the state goes on one side, the possible actions on the other, and a decision is a similarity score between them. So I kept that shape and rebuilt everything around it for this one task:

  • A frozen Qwen3-Embedding-0.6B encoder, using a pinned 8-bit MLX conversion.
  • Two new projection heads, 1024 → 128 → 64, with 278,912 trainable parameters in total. Nothing else is trained.
  • One question over nine combined rotation and thrust actions, instead of two questions. The nine action embeddings are computed once, when the model loads.
  • A shorter observation of 63–69 tokens.

The training data comes from guidance flights on seeds 1000–1031, with validation on 2000–2007. Embedding 3,159 examples and training the heads took 42 seconds on the Mac. Before training, I measured the untouched encoder: raw similarity picked the right action 80.4% of the time. After training the heads, it got every validation example right. That is the loop from above in miniature: measure, train, measure again.

One bench note, because it was not obvious: batching with left padding produced non-finite values in the MLX runtime's attention. Switching to right padding and pooling the last real token fixed it. Batched and single embeddings now agree to a cosine similarity of 0.99996.

Now the best part. Here is the same set of 96 flight states across every deployment I have tried:

 Median msp95 msAgrees with guidance
CLM, released heads + Qwen3-8B Q8869.001,378.0715 / 96
Public Laya, MLX FP1618.5037.3953 / 96
Lunar-trained Laya, MLX Q626.5639.4796 / 96
Small CLM, MLX 8-bit19.7923.9596 / 96

That is a 43.9× lower median than the 8B deployment, and every decision fits inside the 200 ms control interval. On nine flights it had never seen, it landed all nine without a single override, and all 3,351 decisions matched guidance.

I want to be careful about why it got faster, because it is tempting to credit the training. Training made the decisions accurate. It did not make the encoder do less work. The speed came from a smaller encoder, a shorter prompt, one question instead of two, and running in-process with no HTTP. I did not measure how much each one contributed.

Then I put the pads on hills

Replays are nice, but I wanted to poke at it. So the pads moved onto three separate hilltops at 70, 140 and 210 meters, and the replay page became a small interactive app. You click the sky to place the lander, turn a tilt slider all the way around (upside down included), pick a pad and press Launch. A small Python server flies the chosen pilot and streams each decision to the browser as it happens, so playback starts right away.

The hills broke my guidance controller first. It used to hold a fixed height above the target pad, which is a fine rule until a 210-meter hill sits between you and a 70-meter pad. Now it holds 120 meters above the highest terrain between the lander and the pad, climbs before moving sideways when it starts below a ridge, and refuses to fire the engine past 60° of tilt, where thrust pushes sideways or down. It rights itself first. I checked it on a grid of 3,348 valid starts across every pad and tilt, and all of them landed.

And guess what? The small CLM needed no retraining for any of this. Unchanged, it landed nine held-out seeds and 18 custom starts on the new terrain, including the inverted and far-away ones. Every decision matched guidance.

What this does and does not show

Same caveat as the Laya posts, and it matters just as much here. The prompt contains the guidance controller's requested rotation and thrust. The model reads “Requested rotation: left. Requested thrust: half.” and picks the matching action out of nine. Those perfect scores measure how reliably and quickly it follows a request, which is exactly the property I want from a bounded decision, but it is not the model flying on its own.

The comparison table is not a controlled architecture test either. The prompts, the number of questions and the runtimes all differ, and the table says nothing about CLM at data-center scale, where its authors make their speed claims against Jev. The measurements in the repository were taken on the original terrain, where the pads sat in valleys. They are kept as a record of that setup.

Try it

The baseline pilot needs nothing but Python, so you can place the lander yourself in about a minute:

git clone https://github.com/mraad/lunar-clm
cd lunar-clm
python3 -m lunar_clm.serve --pilot baseline

Then open http://127.0.0.1:8000/. The small CLM pilot needs the MLX setup described in the README, and it runs in-process with no servers of its own.

Next, I want to take the requested labels out of the prompt and see what is left. The measurement will tell me exactly which states it gets wrong, and those states become the next training set. That is the self-improvement loop I keep talking about, closed on a model that is finally fast enough to fly inside it. It is future work, not something this post demonstrates. The source code, training details and the full 8B comparison are on GitHub under Apache 2.0. Put the lander upside down over the tallest hill and tell me where it breaks. More to come. :-)

Comments

Popular posts from this blog

How to use Esri Flex API on Android and iPhone

Weighted Map Feature Clustering With Attributes

Augmented Reality, iPhone and ArcGIS Server