Posts

Showing posts from September, 2026

Lunar CLM: From 869 ms to 20 ms per Decision on a Mac

Image
In my last lander post I spent a few paragraphs on Jev , the System One decision model TypeSafe AI opened to early access this month. You hand it a block of state and typed questions, and it returns a probability for every choice. No text to parse. I also made a claim I had not tested: my lander talks to its model through one call, agent.predict(prompt, QUESTIONS) , so swapping in a different decision model should be an adapter, not a redesign. Then I read about Contrastive Language Models (CLM). It takes the same contract and builds it in the open: a block of state in, typed questions, a probability per choice out, with source code and released 8B heads you can run yourself. Its server even answers on an endpoint called /v1/systemone . Jev I can only read about. CLM I could actually test. I had to find out. The adapter part turned out to be true. It was everything after that which kept me busy. The small CLM I trained on my Mac, with no overrides. The lander starts upside do...

The Trainee Got Better. It Still Cannot Fly Alone.

Image
I ended The Engineer and the Trainee with a plan: roll out the shielded pilot, relabel every visited state with MPC, retrain, and see whether the trainee could finally fly on its own. That is still the plan. But before getting there I did something lazier, and the result was more interesting than I expected. I also fixed something that had been bothering me every time I watched a flight. The same recorded flight as last time, re-run with a smoother controller and a retrained model. The engine loses 60% of its thrust at ten seconds, the estimate falls to 40%, and the shield now steps in on about one command in eight instead of one in three. The lander had the shakes Watching a descent, the ship twitched. Not a rendering artifact — the controller really was changing its mind. I measured it: over 30 flights the beam search picked a different command on 77.7% of decisions. Every fifth of a second it re-derived the whole plan from scratch, and when two commands were nearly tie...

The Engineer and the Trainee: Adaptive MPC Guiding and Shielding Laya

Image
Two of my recent lander experiments kept nagging at me. In One Lander, Three Controllers , an adaptive MPC learned its engine during the flight and landed through a power loss with no training data at all. In Lunar Laya , a small decision model on Apple MLX turned a text description of the flight into structured choices, but the text carried hints from a fixed feedback law that assumed a healthy engine. Cut the engine power and that model crashed every time, faithfully following bad advice. So I asked the obvious question: what happens if the engineer writes the hints for the trainee? The result is Lunar MPC + Laya . An actual recorded flight. Laya proposes every command from telemetry alone; adaptive MPC vetoes about one in three. The engine loses 60% of its thrust at ten seconds, the estimate drops to 40%, and the lander touches down at seventy. The engineer and the trainee The game is the same arcade-style lander from Lunar Laya: three pads on bumpy terrain, a tilt control, a...

Lunar Laya: From State and Text to Decisions on a Mac

Image
I keep coming back to Lunar Lander. It gives me a very visible way to ask a simple question: what does a model actually decide? Turn too much, burn too late, or miss the pad, and the result is right there on the screen. After exploring learning and control approaches to this problem, I wanted to try another angle: give a small model the current state and a text instruction, then ask it for a decision that the application can execute and measure. That experiment is now public as Lunar Laya , an Atari-inspired lander using Laya with local inference on Apple MLX. I believe the input should be a state and a text, even a structured text. The output should always be structured and very deterministic, so it can be controlled and measured. This implementation is a representation of that idea. The fact that the model can be small enough to run locally in an MLX environment is proof that it can be done for this kind of bounded decision task. An actual trained-model flight using FP16 MLX ...

Nano FEL: GPT-2 and nanoGPT on the North Sea

After LoRA on a 4B model and a 1.2B model on the Mac , I kept asking myself a simpler question. How small can the model be, and how plain can the training loop be, before this North Sea task stops working? So I went back to the basics: GPT-2 medium, 355 million parameters, fully fine-tuned with Andrej Karpathy's nanoGPT . No adapters, no chat template, no grammar, no catalog in the prompt. Just a question in, a JSON query out. That is nanofel . The task is the same FELN translation: a question about wells, pipelines, and discoveries becomes a plan with layers , where , and relations . The twist here is the data format. nanoGPT trains on one flat stream of tokens and samples random windows from it, so each training pair is simply: Q: Which gas wells in Norway are within 2 kilometers of an oil pipeline? A: {"layers":["Wells","Pipelines"], "where":["content_type = cast(2 as SMALLINT) and (country = 'NO')","Pipelin...

FELN Qwen: Train on RTX, Run on Jetson Orin

Image
After teaching a small model to query the North Sea , I wanted to take the experiment onto a Jetson. Could I train on a powerful workstation, then move the model, application, and spatial database onto an NVIDIA Jetson Orin? And could it produce the expected FELN JSON accurately enough to be useful? That is the experiment behind feln-qwen . The job is specific: translate a question about wells, pipelines, and discoveries into FELN. The JSON identifies the layers, their attribute filters, and the spatial relationships between them. Layer order matters. Finding pipelines near discoveries is a different request from finding discoveries near pipelines! A huge thank you to Geodata for the beautiful NorthSea data . It gave this experiment a rich geospatial vocabulary and meaningful relationships to work with. I deeply appreciate the work and care behind this dataset. I chose Qwen3.5-9B and first verified stock-model inference on the Orin. The target machine had to support the model b...

FELN Liquid: Fine-Tuning a GIS Model on a Mac

I have been experimenting with small language models that translate everyday questions into structured GIS queries. With FELN Liquid , I wanted to explore the training side on the Mac itself. Could I take a compact model, teach it the conventions of a North Sea catalog, and complete the experiment locally? The first complete run used LiquidAI's LFM2.5-1.2B-Instruct with LoRA through MLX on an Apple M4 Max. Training took 48 minutes and 37 seconds. Including checkpoint evaluation and the final test, the workflow finished in just under an hour. That is a practical turnaround for trying an idea and inspecting what it learned. The task is familiar from the other FELN projects: turn a request into a JSON plan containing layers , where , and relations . A compact Layers.json catalog accompanies each request, supplying the names, types, aliases, codes, and hints that make the data meaningful. For example, “Show all wells with water depth > 350 meters” should produce: { ...

FELN WebGPU: A Language Model, a Database, and a Map in One Browser Tab

Image
I have been exploring how small language models can turn a plain-English question into a useful GIS query. With FELN LoRA, that meant training a model on a GPU and running it locally on my Mac. For feln-webgpu , I wanted to bring the language model, the database, and the map into the web application itself. Type “Give me all oil discoveries,” and the browser generates a query plan, compiles it to SQL, executes it, and draws the results. The example in the repository finds 562 discoveries, with 200 displayed. We can inspect the plan and SQL beside the map. I like being able to see what the model actually asked the database! The question, query plan, SQL, and resulting features in one browser tab. The pieces are a fine-tuned LiquidAI LFM2-350M model, Transformers.js running on WebGPU, DuckDB-Wasm with its spatial extension, and the ArcGIS Maps SDK for JavaScript. Once the assets have loaded, query generation and database execution happen on the user's machine. There is ...

FELN LoRA: Teaching a Small Model to Query the North Sea

In the FELN RAG experiment , I used a few relevant examples to help a language model translate a question into a spatial query. That got me interested in taking the next step: teaching a small model to do this particular job, then running it locally. I believe small language models have a big role to play in GIS. We have plenty of tasks with a well-defined vocabulary, a known catalog, and a very specific output. Turning a North Sea question into a query over wells, pipelines, and discoveries is one of them. That is the experiment behind feln-lora . The model is NVIDIA's Nemotron-3-Nano-4B, fine-tuned using LoRA and QLoRA through NeMo AutoModel . Training happens on an NVIDIA GPU. The resulting model runs on my Mac through llama.cpp and Metal, with the question staying on the machine. Now the best part: the latest local models produce a query in about 1.2 seconds at the warm median. That is a useful place to start for an interactive GIS workflow :-) Consider this question fr...

FELN RAG: Five Examples and a Spatial Query

Image
Ask a GIS person to find oil wells within five kilometers of gas pipelines and they can start breaking the request into pieces. Which layer contains the wells? How is “oil” encoded? Which pipelines carry gas? What spatial operation connects the two? A language model needs those same details. “Oil” might be a numeric subtype. A country might be stored as a two-letter code. And choosing intersects when the question calls for contains can produce perfectly respectable JSON that asks the wrong question. That is the problem behind feln-rag : give the model a description of the data and a handful of relevant worked examples, then ask it to translate a new request into a structured spatial query. The interesting part is how little machinery the retrieval needs. Start with the query we want to produce The output is FELN, a JSON representation with three main pieces: layers : the primary layer to return, followed by any spatial filter layers. where : one SQL filter per layer, in the s...

FELN: Find Existing Location with N Layers

In the layers-json post, I argued that the model needs to know what our data means before it can write a useful query. A field called STATUS with a value of 3 is not self-explanatory, and the aliases, domains, and subtypes already sitting in an ArcGIS Pro project are the cheapest explanation we will ever get. That post ended with a catalog, a Layers.json file, and a promise to do something with it. This is the something. FELN stands for Find Existing Location with N layers , and it is a small Python library and CLI that turns a plain-language request into a structured spatial query, compiles that query to DuckDB SQL, and, the part I actually care about, measures how close one query is to another. The shape of a question Most of the spatial questions people ask me over the years have the same skeleton. Find the things in one layer, optionally filtered, that stand in some spatial relationship to the things in another layer, also optionally filter...

Layers JSON: Giving LLMs the GIS Context They Need

When I started writing about GenAI and geospatial analysis , the attraction was being able to ask questions in plain language and let the tools do the spatial work. There is a very practical detail behind that interaction, though: the model needs to know what our data means. Take a field called STATUS with a value of 3 . Is that an active well? An abandoned one? A record waiting for review? A model can write perfectly valid SQL around that number and still give us the wrong answer. The database will happily execute it, too :-) That is the problem I want to address with layers-json : give applications access to the knowledge already captured in an ArcGIS Pro project and its supporting data sources. The map already carries a lot of the explanation If you have spent time giving fields useful aliases, defining domains, and organizing layers in Pro, you have already done some of the hard work. A field alias turns a storage name into something a person recognizes. A coded-value doma...

Nano Face: A Mac Camera, a Jetson, and One USB Cable

Image
Let me get the obvious out of the way first. Face detection on a Jetson is not new. NVIDIA ships samples for it, half the tutorials on the internet do it, and I am fairly sure somebody has done it on a Raspberry Pi wearing a HAT. So no, this is not a post about discovering face detection. This is a post about my two Jetsons, an Orin Nano Super and an AGX Orin, that live headless on a shelf with no monitor, no keyboard, and no camera plugged into either of them. I had just finished moving the Nano to JetPack 7.2.1 and pruning both boards down to a lean server install (no desktop, no Docker, nothing starting at boot), and I wanted a small, honest, end-to-end test that they were alive and doing real work. A face detector is a good smoke test. It either draws a box on your nose or it does not. The twist, and the part I actually cared about, is where the camera lives. It is the one in my MacBook Pro. The Jetson never sees a camera device. Frames leave the Mac...

ArcGIS Pro, Claude Code, and a Loopback Bridge

Continuing the GenAI-with-a-GeoSpatial-twist thread I started a while back. Back then, I had a language model reason about geospatial logic. This time, I wanted it to actually do the work—on my open project, inside my running ArcGIS Pro session, while I watch. The result is ProCowork , an experimental native ArcGIS Pro add-in that embeds the Claude Code engine in a dockable chat panel. I can ask it to list layers, add and calculate fields, select features, run a buffer, or write a more specialized ArcPy script. Claude generates the code, the add-in runs it against the live project, and the code and results come back into the same panel. For example, I can type: “List the layers in the current map.” “Add a DOUBLE field POP_DEN to Parcels and set it to POP / AREASQMI.” “Select parcels where POP_DEN > 5000 and zoom to them.” “Buffer Roads by 100 meters and add the result to the map.” The panel is useful, but the interesting part is ...