Overview
MetaMedium is a way of working with AI by drawing. You sketch on a canvas and the canvas reads what you drew — a circle is a circle, two boxes side by side are a row, a word beside a shape is its name — and offers its reading back where you can see it and argue with it. Name a pattern and the canvas learns your vocabulary. A model can join, reading with you and drawing back. The drawing stays yours.
The idea underneath: AI as a meta-word, a new part of language that turns rough marks into meaning from their context. Drawing becomes something a person and a machine can do together, with a seven-year-old's sketch and an engineer's diagram on the same canvas, each read at its own level.
The loop it describes is replayed just below — then the argument for it.
Running the Loop
This is the canonical loop, run once through the engine and recorded as the events it produced. What you step through below is not a video: it is the engine replaying its own log in your browser, and the inspector on the right holds every reading at every step. Draw on it at any point and the session continues with your marks.
The Problem
Dead Drawing on a Living Medium
Today's drawing tools, from Illustrator to OneNote, treat drawings as pixels or vectors, not as meaning. Illustrator knows a circle's geometry but not what it stands for. OneNote captures your marks and never interprets them. Procreate's strokes stay forever inert.
Even Figma and Miro treat diagrams as layout, not computation. You can draw a flowchart, but the arrows do not flow. You can sketch a state machine, but it does not run. The computer records what you drew without understanding what you meant.
Drawing, our most natural way of working an idea out, stays inert on the screen. We have given computers eyes, ears and a voice, but not a way to think alongside us on paper.
The Communication Bottleneck
Language models can now reason, write and hold a conversation. We reach them through a text box — like sharing a piece of music by describing it.
The limit is the channel, not the model. We think in pictures, space and gesture, and we have a text box to say it through. A designer describing a sketch in words, a child turning play into a query — something is lost every time.
This is about more than convenience. A text box favours people who already think in text. A child who thinks in pictures, a craftsperson who thinks with their hands, an elder who thinks in stories — each is left at the edge of what these machines could do with them.
"In a few years, men will be able to communicate more effectively through a machine than face to face." — J.C.R. Licklider & Robert Taylor, "The Computer as a Communication Device," 1968
It has come true for text. We have richer languages already — drawing, gesture, arrangement, annotation, showing — and no interface that speaks them.
The Vision: As We May Sketch
In 1945, Vannevar Bush imagined the Memex—a device for extending human memory and enabling associative thinking. He asked: As we may think, how might machines augment the trails of connection that constitute human understanding?
We ask the parallel question. As we may sketch: how might machines carry the visual, spatial, gestural thinking that so much of human thought is made of — especially in children, who draw before they write and think in pictures before they think in propositions?
Anything digitized has become an abstraction, so let's embrace it. When I draw into a computer with the flourish of my hand, we can take it beyond pixels, beyond even vectors, toward universal mapping attempts, toward a truly metamedium. — John Hanacek, "As We May Sketch," Georgetown CCT Masters Thesis 2016
A curve drawn by hand is a chance to try fitting a function to the line. A function is just waiting to become metaphorical graphics. Digital ink will move beyond "networked paper" to become a magical plane where computer vision partners with the human hand functioning as an interactive external imagination. From sketch to code, from code to sketch—no longer a pipeline but rather a constellation of possibilities, an ever-expanding network of opportunities to map expressiveness and flow to logic and math directly.
Dancing Without Music
"Imagine that children were forced to spend an hour a day drawing dance steps on squared paper and had to pass tests in these 'dance facts' before they were allowed to dance physically. Would we not expect the world to be full of 'dancophobes'?" — Seymour Papert, Mindstorms, 1980
Papert's question cuts to the heart of how we teach abstraction. We have raised generations of "mathphobes": people who believe they are bad at math while routinely using logical reasoning to fix computers, build furniture and run businesses.
The problem is not aptitude. We ask people to dance without music: to manipulate symbols divorced from meaning, to learn the steps before they feel the rhythm. A child with an intuitive sense for assembling things in space may never connect school geometry with building.
Give that same person constant feedback, immediate results, symbol connected to meaning through direct manipulation, and they discover they were never unable to do math. They were never shown the connection.
A Medium for Children
"The child is a 'verb' rather than a 'noun', an actor rather than an object... We would like to hook into his current modes of thought in order to influence him rather than just trying to replace his model with one of our own." — Alan Kay, "A Personal Computer for Children of All Ages," 1972
Children already think computationally: in systems, in cause and effect, in what happens if. They do not need to learn to code first. They need an interface that meets them where they are — drawing, playing, exploring. MetaMedium is built for that: a canvas where a child's "bouncy house" already has the beginnings of structural engineering in it, and a doodled spiral is a way into mathematics.
The Lineage
MetaMedium stands on decades of work in sketch interfaces and computational media. The lineage shows what is new here and what is borrowed.
- From Visions: Dynabook's metamedium concept + Victor's directness principle
- From Recognition: Sketch-editing games' negotiation paradigm
- From Intelligence: LLM interpretation + probabilistic reasoning
The Thesis: AI as Meta-Word
Closing the Triadic Loop
Today's interfaces connect language to computation and leave meaning outside. We write; the machine executes; something comes back. What it meant lives only in our heads, before and after.
Traditional interfaces flow one way: we write, machines execute, outputs return. Meaning remains external.
When a mark can mean several things, and the system holds those readings and refines them with you, meaning becomes part of the loop instead of something that happens off-screen.
AI as Meta-Word
Writing externalised memory: a thought could be kept and picked up again later. AI can externalise interpretation — the work of making meaning from marks in context. That is what a meta-word is for.
What Is a Meta-Word?
A meta-word is not a word about words, like "noun" or "verb". It is a word that changes other words. Read "bounce" beside a spring and it knows what a spring does when it bounces, and can make it happen.
In ordinary communication people trade signs — words, gestures, marks — and each rebuilds the meaning for themselves. MetaMedium lets the AI take part in that rebuilding. It:
- holds several readings of a mark at once
- keeps them open until context settles it
- learns your vocabulary
- moves between sketch, equation, code and text
- Less a tool than a part of speech.
"Thanks to a mapping, full-fledged meaning can suddenly appear in a spot where it was entirely unsuspected." — Douglas Hofstadter, I Am a Strange Loop, 2007
Thinking as Conceptual Blending
What is human thought anyway? Gilles Fauconnier and Mark Turner attempted to answer this with their theory of conceptual blending, building on Lakoff and Johnson's work on how metaphor structures understanding. Consider a riddle:
A Buddhist monk begins at dawn walking up a mountain, reaches the top at sunset. After several days, he walks back down, starting at dawn and arriving at sunset. Is there a place on the path he occupies at the same hour on both journeys?
The answer becomes obvious the moment you visualize two monks walking the path simultaneously—one going up, one going down. They must meet somewhere. But this visualization requires what Fauconnier and Turner call an "integration network"—a blended mental space where separate inputs combine to reveal emergent structure. I have animated their central figure illustrating the blending space as a diagram.
The MetaMedium is a system for building integration networks on a canvas. When you draw a diagram, you are setting up mental spaces. When you connect elements with arrows or proximity, you are creating cross-space mappings. When the AI interprets your marks and offers possibilities, it is helping locate shared structures. Diagrammatic thinking externalizes the blending process—making it visible, manipulable, shareable. Two people looking at the same diagram can point to the same conceptual space.
Tools vs. Medium
Alan Kay's Dynabook vision asked: "What is the carrying capacity for ideas of the computer?" His answer: the computer is a metamedium—it can simulate any existing media and also be the basis of media that can't exist without the computer. But Kay made a crucial distinction:
"What then is a personal computer? One would hope that it would be both a medium for containing and expressing arbitrary symbolic notions, and also a collection of useful tools for manipulating these structures." — Alan Kay, "A Personal Computer for Children of All Ages," 1972
Most AI interfaces treat the model as a tool: something you call, ask, command. MetaMedium puts it in the medium. You are not using the computer to sketch; you are sketching in a material that can read, and the AI is the part of the material that understands.
Everything interactive in this paper is a recording from the same engine. The surface itself — live, drawable, yours to try — is in Current Development.
The Framework
Core Principles
Space Is Semantic
Spatial relationships carry meaning. Near means related; a line means a directed relation. Position, proximity and connection mean something on their own.
Place two circles close together; the system infers "related." Draw one inside another; it understands "containment." Position creates meaning without words.
Built. Nearness, insideness, alignment and direction are measured as ratios of the marks' own size, carry a strength, and are what the palette offers from and the model is briefed with.
Annotation Becomes Execution
Write "make this bounce" beside a spring and the note is an instruction. Draw an arrow from input to output and you have defined a flow. To describe is to instruct.
Write "3x" next to a line; it becomes three lines. Write "wiggle" near a shape; it animates. The annotation is the program.
Partly. A word written beside a shape is read and offered as its name; a prompt on a circled group builds a page in place, and ink on that page addresses the region under it. “3x” and “wiggle” are still vision.
Ambiguity Is a Feature
A rough sketch is understood as rough. The system holds several interpretations and refines them as context accumulates, the way people tolerate ambiguity and resolve it over time.
Your rough oval might be a face, an egg, or a zero. The system holds all three until you add two dots — then it settles on "face." Deciding too early ends the exploration.
Built. A mark holds every reading that qualifies, ranked by measured confidence. A pentagon is rectangle and circle at once, and is redrawn clean as neither.
Bidirectional Learning
The system learns your vocabulary and conventions into a cognitive lens, and teaches you back by surfacing patterns and suggesting relationships. Your notation becomes something the canvas can act on, and its readings become something you can see.
Draw "recursion" shorthand repeatedly; the system learns it. Later, it suggests this mark when detecting recursive patterns—teaching you to see what it sees. Your notation becomes shared language.
Partly. Draw your command mark five times and it becomes yours; name a group and the next one like it is recognized. Lenses beyond that are not built.
No Mode Switching
Following Larry Tesler's "no modes is good modes": you are always just working, drawing, annotating, refining. Interpretation appears when needed and fades when not.
Draw a shape. Write near it. Adjust with gestures. Watch it execute. All the same canvas, all the same moment. The interface disappears into the work.
Built. Selection, command and erase are marks: a loop is a lasso, your mark across it summons, a scratch erases what it crosses. There is no mode to be in.
Observable Reasoning
Uncertainty is visible. When several interpretations are held you see them, not a single guess, and they stay present until context or your choice resolves them.
Your rough mark triggers three possible interpretations shown as faint ghosts; tap one to commit, or keep drawing to refine. You see the system thinking.
Built. Every reading names the measurement it rests on; a confident one ghosts its clean form under the ink; an inspector walks any mark from ink to shape to role to code.
The Negotiation Paradigm
Ribeiro and Igarashi's "Sketch-Editing Games" (UIST 2012) introduced a negotiation paradigm where user and machine take turns refining interpretation. The user sketches; the machine recognizes and offers interpretations ("bottle?"). The user refines ("no, more like a mug"). The machine updates. Understanding emerges through iterative exchange.
Their key insight: sketch recognition improves dramatically when reframed as a game rather than a classification problem. The machine maintains a "possibility graph"—a network of possible interpretations and the transformations that would select among them. The user's next stroke navigates this graph, collapsing some possibilities and opening others.
MetaMedium makes the possibility graph learnable, accumulating your patterns over time, and adds annotation as another way to navigate it. A misreading is information. Each correction adds to a shared vocabulary.
The Semiotic Foundation
Charles Sanders Peirce described meaning-making in a way that fits human–AI work well. In his model, meaning comes from the relation between the sign (the form), the object (what it stands for), and the interpretant (the meaning made in the interpreter's mind).
Peirce's insight — that meaning is rebuilt, not transferred — is the Language ↔ Computation ↔ Meaning loop from earlier. In MetaMedium the AI takes the interpretant's seat: it holds possible meanings and refines them with you, instead of executing a command.
Cognitive Lenses
As patterns accumulate, they form "cognitive lenses"—personalized interpretation frameworks that shape how the system reads new marks. A physicist's lens recognizes force diagrams; an architect's lens sees load-bearing structures; a musician's lens interprets spatial arrangements as rhythm.
Lenses can be shared. A research group might build one for its notation; a classroom might inherit one from whoever designed the course. A lens is a way of seeing, and it can be handed on.
Current Development
A working engine and a reference surface accompany this paper. You draw on an infinite canvas; the canvas reads what you drew; a model can join the reading. Every mark climbs three rungs, each a closed vocabulary you can inspect:
- Shape: what the stroke is. Line, arc, triangle, rectangle, circle, arrow, writing, dot. Every reading is measured from the ink and carries its reason; several are held at once. A confident one is offered back as its clean form, drawn over the ink, never in place of it.
- Diagram: what the mark plays. Container, node, edge, label, annotation, placed from measured relations, so a drawing has a genre: boxes tiling a space are a page, nodes joined by edges are a graph.
- Code: what it becomes. A page compiles to flexbox that reflows, your ink still outlining its elements; a graph keeps its positions and arrows. The engine owns the structure it measured; a model writes only the content, briefed by the drawing.
Selection and command are marks too. Circle a group, cross it with a command mark you taught by drawing it five times, and a palette offers what those marks could become, starting with what needs no model. Handwriting beside a shape is read by a model that can see and offered as the shape's name. The model can draw back, in the shapes the canvas can read, its marks held in its name beside yours.
The loop the product is built around, recorded once against a local eight-billion-parameter model and replayed here by the engine: four boxes become a page inside the ink, and a mark drawn on the running page changes only the region it lands on.
And the surface itself, live. Everything here works offline; join a local model in the full page to build, read handwriting, or have it draw.
Demo & Source
Everything above works offline; joining a local model (Ollama or LM Studio) or a hosted one by key adds the reading, building, handwriting and drawing that need one. Development continues in the repository.
Reference surface: jjh111.github.io/MetaMedium/Demos/session-engine.html
Earlier prototype (2025): doodle2-canvas.html — heuristic recognition, a learned library, and the geometry of what you drew read out as maths
GitHub: github.com/jjh111/MetaMedium
License: GPL — the MetaMedium framework is open source.
Development Roadmap
- Built: the engine and the three rungs; selection, command and erase as marks; living artifacts that ink can address; local and hosted models as participants; handwriting; the model drawing back.
- Next: words from printed letters; the model proposing library entries for the human to bless; recall by meaning; this surface as the flagship.
- Then: multi-user canvases; cognitive lens export and import; a model outside the browser taking part through the same channel.
Open Questions
The MetaMedium framework raises questions that can only be answered through building and testing:
- How do we design for "interpretive ambiguity" without creating confusion? What's the right balance between holding possibilities open and committing to interpretation?
- What are the limits of gesture vocabulary before cognitive overhead exceeds benefit? How many "words" can a visual language productively contain?
- How can "cognitive lenses" be effectively shared between users? What's lost in translation when one person's way of seeing meets another's?
- What comes out of drawing with a model that neither would make alone?
- Can visible reasoning on a shared canvas measurably improve AI alignment outcomes? This is empirically testable.
Limitations and Challenges
- Recognition is bounded, not solved: the engine reads eight shapes and one cursive stroke per word; everything past that vocabulary is held as "art" until a model or the human says otherwise. The negotiation paradigm turns each miss into a correction, but too many corrections and people stop drawing.
- Cognitive load: Every gestural vocabulary is a language to learn. There's a real risk that MetaMedium becomes its own expertise barrier, replacing "learn to code" with "learn our gestures." Keeping the learning curve gentle while enabling power is a design challenge.
- Privacy of patterns: If the canvas learns from your drawing, those patterns become data. A cognitive lens is intimate — it encodes how you think. Privacy has to be built in from the start.
- Over-automation risk: "The system guessed wrong and did something I didn't want" is a real failure mode. Undo must be instant and obvious. Interpretations must be inspectable before they execute. The user must remain in control.
- Evaluation difficulty: How do we measure success? Traditional usability metrics may not capture "quality of thought." New evaluation frameworks are needed.
Abstraction Management and Learning Dynamics
A canvas that learns creates its own problems. Vocabulary accumulates and nothing forgets, so old notations compete with new ones. Worse, a system well fitted to your previous way of thinking may resist your attempts to evolve, correcting you back toward familiar patterns just when you are trying to break a frame. Possible mitigations include explicit unlearn gestures, decay with different rates for core and peripheral vocabulary, versioned lens snapshots, and treating systematic deviation as a signal. The deeper question remains: is the canvas a memory of what you have done, or a partner in what you are becoming?
Gallery
Scenarios
The framework becomes concrete through scenarios. Each demonstrates specific principles in action.
Visual Learning
Principles: Space is semantic, Canvas learns, Bidirectional representation
Visual thinker draws parabolas; system connects spatial intuition to formal equations bidirectionally. Discovers he understood calculus all along—just needed symbols connected to drawings. Read Story
Asymmetric Collaboration
Principles: Interpretive ambiguity, Cognitive lenses, Negotiation
Seven-year-old's playful "bouncy bridge" sketch becomes engineering student's seismic dampening simulation. Canvas holds both interpretations—intuitive play and rigorous analysis—without translation. Read Story
Rapid Prototyping
Principles: Annotation becomes execution, Negotiation paradigm
Non-programmer sketches water tanks, annotates flow logic. System generates simulation, asks clarifying questions, updates as she refines. Continuous negotiation from rough idea to working prototype—no code. Read Story
Scientific Collaboration
Principles: Space is semantic, Annotation becomes execution, Shared lenses
Researchers sketch faster than formal notation allows. Spatial annotations like "defect here?" trigger simulations. Shared research lens interprets their shorthand. Read Story
The Future
External Imagination
The person steers: they know which possibilities matter. The model holds many at once and keeps them in view. It is an external imagination — the wind under your own thinking, carrying it further than it would go alone, while you still do the flying.
"Artificial" is the wrong word; it says fake, lesser. Human intelligence is embodied, mortal, shaped by living. Machine intelligence is computational, distributed, shaped by training. Both are real, and the question is what they can do together that neither does alone.
Alignment Through Communication
Most approaches to alignment focus on control: rules, guardrails, constraints, on the assumption that the machine's goals might diverge from ours. The MetaMedium proposes an alternative: enrich the medium between us so that coordination happens through communication. The richer the shared vocabulary, the better we can align our understanding.
That is how we coordinate with each other: not by controlling one another's thoughts but by sharing a medium rich enough to work things out in. The canvas can be that medium, with both sides' reasoning on it where both can see.
Beyond 2D: Navigating Conceptual Space
The current framework treats diagrams as 2D arrangements, but diagrams are projections of higher-dimensional conceptual space. Future development could explore navigating the space a diagram lives in—not just the diagram itself. Three-dimensional visualization would give canvas elements depth: z-axis as semantic distance, uncertainty, or abstraction level. Four-dimensional (temporal) visualization would make the evolution of understanding navigable—scrub through versions, see where insight branched, experience collaborative history as visible geology.
Most speculatively: latent space rendering. AI models maintain high-dimensional embedding spaces that encode meaning. What if the canvas could project these spaces, letting users see where their current sketch sits relative to possible interpretations? The "possibility graph" becomes navigable terrain; your marks become waypoints through semantic space.
The Deeper Vision
There is a version of this vision that goes beyond interface. Today's computing is an archaeological site: layer upon layer of abstraction, each solving problems created by the layer below, each adding distance from what the machine does. Bret Victor's "Future of Programming" reminds us that direct manipulation, visual programming and goal-directed systems were explored in the 1960s and then buried under commercial code and decisions no one remembers. Now AI arrives, and we add more layers.
The deeper vision goes the other way: the computer knowing what it can do and doing only as much as it needs to, every operation justified, every layer earning its existence. AI could be the tool for this, reading the whole stack and finding the essential operations under the accretion. On the canvas, interpreting a sketch could mean "generate Python", or it could mean "this is a constraint problem; here is how it maps closer to the metal." The diagram negotiates the level of abstraction the thought requires.
The metamedium dream waits beneath the APIs and the bloat, patient, ready to be excavated. The tools to dig are finally arriving.
Conclusion
The limit has been the channel, not the model. The MetaMedium is the surface where a person brings their whole way of thinking to a machine that can read: a mark is a proposal, the machine's reading is visible and arguable, and the two of them are building the same drawing. What it does not yet do is listed above as plainly as what it does. Development continues. It's time to bring the computer to life at the depth of mind with the speed and intuitive action of our hands.