Layers JSON: Giving LLMs the GIS Context They Need

When I started writing about GenAI and geospatial analysis, the attraction was being able to ask questions in plain language and let the tools do the spatial work. There is a very practical detail behind that interaction, though: the model needs to know what our data means.

Take a field called STATUS with a value of 3. Is that an active well? An abandoned one? A record waiting for review? A model can write perfectly valid SQL around that number and still give us the wrong answer. The database will happily execute it, too :-)

That is the problem I want to address with layers-json: give applications access to the knowledge already captured in an ArcGIS Pro project and its supporting data sources.

The map already carries a lot of the explanation

If you have spent time giving fields useful aliases, defining domains, and organizing layers in Pro, you have already done some of the hard work. A field alias turns a storage name into something a person recognizes. A coded-value domain explains what those numbers mean. Subtypes tell us which kind of feature a row represents. Metadata descriptions and sample values fill in more of the picture.

I like the idea of putting that work to use beyond the map. Someone took the trouble to explain the data once; we should be able to carry that explanation into the next application.

The layers-json command reads the project's layer configuration together with its referenced File Geodatabase or PostgreSQL data and writes a Layers.json catalog. It brings together aliases, field types, domains, subtypes, sample values, and query hints. A consuming application can supply the relevant parts to an LLM when translating a question into SQL.

For our hypothetical well-status domain, that means the application can supply the actual mapping between “active” and its stored code. Now the model has something concrete to work with.

From an APRX to JSON, or something we can read

Once the package and its GDAL dependencies are installed, a catalog command looks like this. Replace the example path and map name with your own:

layers-json /data/NorthSea/NorthSea.aprx --map Map -o ./catalog

This writes catalog/Layers.json. The command runs outside ArcGIS Pro, without ArcPy. It needs Python 3.11 or newer, system GDAL, and matching GDAL Python bindings. The README has the installation steps. If you prefer working in Pro, the bundled Layers.pyt toolbox provides a Prepare Metadata tool for the active map.

There is also a format I can open and read directly:

layers-okf /data/NorthSea/NorthSea.aprx --map Map -o ./knowledge

This uses the same reader to produce an OKF knowledge bundle: an index.md and Markdown concept files for the tables. The bundle includes schema tables, decoded domains, query hints, and source references, with structured metadata in the index. It gives us a readable way to inspect the context before handing it to an application.

There are some useful boundaries here. Sampling defaults to at most 20,000 rows per layer and 20 sampled distinct values per field. Columns without sampled or domain values are pruned. I would inspect the result before treating it as the context for a query; it is a catalog built to help querying, and sampling can miss values.

And the actual rows?

The catalog describes the data. To work with the rows in a SQL database, the repository also includes export tools. From an ArcGIS Pro Python environment with the DuckDB extra installed, we can export a project's layers:

layers-duckdb "C:\Projects\NorthSea\NorthSea.aprx" --map Map

That exporter writes a sibling DuckDB database, preserves hidden attributes and GlobalIDs, honors definition queries, and converts geometry to EPSG:4326. Separate GDAL-based Bash loaders handle File Geodatabases going into DuckDB, PostGIS, or Oracle Spatial. The script reference explains their requirements and replacement behavior.

What interests me most is what this could do for smaller language models. Supplying the meaning of a field leaves less for a model to infer. That is the motivation; this project does not yet establish an accuracy improvement through model benchmarks. The application still has to select the context, generate the query, and check the answer.

Try it with a small project whose data you know well. Open the catalog, look at the aliases and coded values, and see whether it contains the explanation you would give a colleague asking about that dataset. That is a useful place to start.

The source is on GitHub. I am interested in which pieces of GIS context make the biggest difference for your queries. Feedback and questions are welcome. More to come!

Comments

Popular posts from this blog

How to use Esri Flex API on Android and iPhone

Augmented Reality, iPhone and ArcGIS Server

Weighted Map Feature Clustering With Attributes