Nano FEL: GPT-2 and nanoGPT on the North Sea
After LoRA on a 4B model and a 1.2B model on the Mac , I kept asking myself a simpler question. How small can the model be, and how plain can the training loop be, before this North Sea task stops working? So I went back to the basics: GPT-2 medium, 355 million parameters, fully fine-tuned with Andrej Karpathy's nanoGPT . No adapters, no chat template, no grammar, no catalog in the prompt. Just a question in, a JSON query out. That is nanofel . The task is the same FELN translation: a question about wells, pipelines, and discoveries becomes a plan with layers , where , and relations . The twist here is the data format. nanoGPT trains on one flat stream of tokens and samples random windows from it, so each training pair is simply: Q: Which gas wells in Norway are within 2 kilometers of an oil pipeline? A: {"layers":["Wells","Pipelines"], "where":["content_type = cast(2 as SMALLINT) and (country = 'NO')","Pipelin...