Full AI Suite for LispE: llama.cpp, tiktoken, MLX and PyTorch

Full AI Suite for LispE: llama.cpp, tiktoken, MLX and PyTorch

I have presented LispE a few times in this forum. LispE is an Open Source version of Lisp, which offers a wide range of features, which are seldom found in other Lisps.
I have always wanted to push LispE beyond a simple niche language, so I have implemented 4 new libraries:

  1. lispe_tiktoken (Openai tokenizer)
  2. lispe_gguf (encapsulation of llama.cpp)
  3. lispe_mlx (Mac OS's own ML library encapsulation)
  4. lispe_torch (An encapsulation of torch::tensor and SentencePiece, based on PyTorch internal C++ library)

I provide the full binaries of these libraries only for Mac OS (see Mac Binaries).

What is really interesting is that the performance is usually better and faster than Python. For instance, I provide a program to fine-tune a model with a LoRA adapter, and the performance on my Mac is 35% faster than the comparable Python program.

It is possible to load a HuggingFace model, to load its tokenizer and to execute inferences directly in LispE. You can also load GGUF models (the llama.cpp format) and run inference directly within LispE. You can download models from Ollama or LM-Studio, which are fully compatible with lispe_gguf.

The MLX library is a full fledged implementation of the MLX set of instructions on Mac OS. I have provided some programs to do inference with specific MLX compiled models. The performance is on par and often better than Python. I usually download the model from LM-Studio, with the MLX flag on.

The whole libraries should compile on Linux, but if you have any problems, feel free to open an issue.

Note: MLX is only available for Mac OS.

Here is an example of how to load and execute a GGUF model:

    ; Test with standard Q8_0 model
    (use 'lispe_gguf)
    
    (println "=== GGUF Test with Qwen2-Math Q8_0 ===\n")
    
    (setq model-path "/Users/user/.lmstudio/models/lmstudio-community/Qwen2-Math-1.5B-Instruct-GGUF/Qwen2-Math-1.5B-Instruct-Q8_0.gguf")
    
    (println "File:" model-path)
    (println "")
    (println "Test 1: Loading model...")
    
    ; Configuration: uses GPU by default (n_gpu_layers=99)
    ; For CPU only, use: {"n_gpu_layers":0}
    (setq model
       (gguf_load model-path
          {"n_ctx":4096
             "cache_type_k":"q8_0"
             "cache_type_v":"q8_0"
          }
       )
    )
    
    ; 2. Generate text only if model is loaded
    (ncheck (not (nullp model))
       (println "ERROR: Model could not be loaded")
       (println "Generating text...")
       (setq prompt "Hello, can you explain what functional programming is?")
       ; Direct generation with text prompt
       (println "\nPrompt:" prompt)
       (println "\nResponse:")
       (setq result (gguf_generate model prompt {"max_tokens":2000 "temperature":0.8 "repeat_penalty":1.2 "repeat_last_n":128}))
       (println)
       (println "-----------------------------------")
       (println (gguf_detokenize model result)))
Why is it different?

One of the first important things to understand is that when you are using Python, most of the underlying libraries are implemented in C++. This is the case for MLX, PyTorch and llama.cpp. Python requires a heavy API to communicate with these libraries, with constant translations between the different data structures. Furthermore, these APIs are usually pretty complex to modify and to transform, which explains why there is a year-long backlog of work at the PyTorch Foundation.

In the case of LispE, the API is incredibly simple and thin, which means that it is possible to tackle a problem either as LispE code or when speed is required at the level of the C++. In other words, LispE provides something unique: a way to implement and handle AI both through the interpreter or through the library.

This is how you define a LispE function and you associate this function with its C++ implementation:

        lisp->extension("deflib gguf_load(filepath (config))",
                        new Lispe_gguf(gguf_action_load_model));

On the one hand, you define the signature of the library function, which you associate with an instance of a C++ object. Once you've understood the trick, it takes about 1/2 hours to implement your own LispE functions. Compared to Python, there is no need to handle the life cycle of the arguments, this is done for you.

        Element* config_elem = lisp->get_variable("config");
        string filepath = lisp->get_variable("filepath")->toString(lisp);

The name of your arguments is the way to get their values on top of the execution stack. In other words, LispE handles the whole life cycle itself, no need for PyDECREF or other horrible macros.

LispE is close to the metal

One of the most striking features of LispE is that it is very close to the metal in the sense that a LispE program is compiled as a tree of C++ instances. Contrary to Python, where the code in the libraries executes outside of the VM, LispE doesn't make any difference between an object created in the interpreter or into a library, they both derive from the Element class and are handled in the same way. You don't need to leave the interpreter to execute code, because the interpreter instances are indistinguishable from the library instances. The result is that LispE is often much faster than Python, while proposing one of the simplest APIs to create libraries around.

What is next?

The lispe_torch library is still a work in progress, for instance MoE is not implemented yet in the forward. In the case of tiktoken, gguf and MLX, the libraries are pretty extensive and should provide the necessary bricks to implement better models.

Discussion 4 comments · 7 points · Claudius · 2026-01-30
Open on Lobsters
Loading the discussion…

Domain filters

Stories from these domains are hidden from every list. Subdomains match too: blocking substack.com also hides danluu.substack.com.

    New collection

    Delete this collection?

    About YAVCHN

    YAVCHN is a reader for Hacker News and Lobsters, with articles and discussions in separate windows or Classic pages.

    Created by Paul Parks and built with PUDL.

    YAVCHN source code on GitHub

    Privacy policy · Terms of use

    Help

    Keyboard

    j / k
    Move down and up the story list. The arrow keys scroll whatever has focus.
    Enter
    Read the marked story in the article reader.
    ]
    Read the next story in the same article-reader applet. Back returns to the previous story.
    p
    Pin or unpin the marked story, which keeps it in Pinned.
    n / N
    Move to the next or previous top-level comment in the window in front.
    c
    Collapse or expand that comment.
    f
    Hide or show the story list.
    Esc
    Close a menu or this help.
    Access key m
    Go to the menu bar. Most browsers take it with Alt on Windows and Linux, and Safari with Control and Option.
    ?
    Show this help.

    Windows

    Each story opens in a window holding its article above its discussion; drag the bar between them to share the room differently. A window can be moved by its title bar, resized from any edge, snapped to a half or a corner by dragging it there, maximised, or minimised to the bar at the foot of the page. Use Window > New reader window to open an empty reader, or Story > Open in new reader window to open another reader for the current article. Docked readers keep their articles when you select another story from the sidebar. Minimized readers can be restored and reused for their site. A window's Next story link reads on down the list in the same window.

    A link in a comment or an article to another Hacker News or Lobsters thread opens that thread in a window too. A link to a single HN comment opens the comment above its replies.

    While a story's window is in front, the Story and Discussion menus in the menu bar hold its commands: pinning, Next story, sorting, collapsing every thread, jumping to the first new comment. Each window also remembers where you were in its article and discussion, so a reload, or Back to a story that Next took you past, finds your place again. Closing a window forgets it.

    The whole arrangement lives in the address, so a bookmark or a shared link brings it back, and Back undoes the last change. Moving between Hacker News, Lobsters, their lists, Pinned and Find changes only the list, and leaves the windows open.

    The list

    The pin at the start of a row keeps the story in Pinned, and the cross at its end hides it. Pinned can be narrowed by words in the title, site or author, by source, and to the stories you haven't opened yet, and ordered by when you pinned them, by points or by comments; the filters are part of the address, so a filtered view can be bookmarked. Scroll past the end of the list to load more. Domain filters, in the View menu, hide every story from a site.

    Collections are named lists of stories. Story > Add to collection files the story in front into one or more of them, and the Collections feed shows them all or one at a time; the menu that chooses collections also creates, renames, and deletes them. A note is your own text on a story. Choose Add note in a story's toolbar to write one; it saves as you type. Rows with a note carry the note mark, and the Notes feed lists every noted story and searches the text of your notes.

    Browsing view

    View > Windowed and View > Classic select the browsing view and save your default in this browser. Window view reuses a reader for each feed. Classic view opens stories and applets as pages. Open as a page is a one-off action that does not change your saved default. Use the Windowed selector to return an article to a window. Direct page links always open as pages.

    Applets

    The Applets menu in the menu bar holds three tools, each a window of its own. Replies to me takes your Hacker News user name and lists the replies to your last thirty comments and stories, checking again every three minutes while it is open, and marking what is new since you last marked them read. Look up a user opens a profile on Hacker News or Lobsters, with their submissions and recent comments, as a commenter's name in any discussion does; the bar at the top of a profile looks up someone else in the same window, and Back returns to the one before. Who is hiring? filters the posts of HN's monthly hiring threads by the words you type.

    They read only what the sites publish to everyone, so none of them asks for a login, and your user name stays in this browser anonymously and is included in retained applet state when signed in.

    Find

    Find takes any link and lists every time it was submitted to Hacker News and Lobsters, so you can read each discussion of it.

    About

    YAVCHN never sees your Hacker News or Lobsters login. The discussion is fetched from each site's public API; to vote or reply, follow the link above the discussion, or the arrow beside a comment, to the source's own site. Without a YAVCHN account, your data stays in this browser. When signed in, pins, collections, notes, blocked domains, and retained reading state are stored with your account and synchronized across devices. Hidden stories and layout stay in this browser. The privacy policy has the details.

    Open source: github.com/paulmooreparks/yavchn. Built with PUDL.