Visual-spatial perception for AI agents

Agents want to see like we do.

Giving AI a sense of shape, distance and how things fit together.

Every AI that sees should see like we do.

sehn = “to see” · say “zane”

The problem

An LLM doesn’t see an image. It reads a list of patches.

Today’s models talk brilliantly about pictures. Yet they stumble on spatial tasks a small child solves at a glance. Try one yourself.

Three tangled lines start at dots A, B and C on the left and end in boxes 1, 2 and 3 on the right. A B C

Roughly what the model gets: separate patches

Which box does line A end in?

Tap the box where line A ends.

A three-year-old simply follows the line with their eyes.

A vision model first cuts the picture into a grid of small patches and turns each one into numbers. A line that runs across many patches is no longer one line, so the model has to guess where it goes.

The same blind spot shows up everywhere: what is in front of what, what connects to what, where there is room to move.

How it works

One sentence, and your agent can see.

sehn.ai is an MCP service. Tell your agent to use the tools at sehn.ai, and it gains a sense of space it did not have before.

  1. 01

    Any agent

    Bring the agent you already use.

    • Claude
    • GPT
    • Gemini
    • Any agent
  2. 02

    The sehn.ai MCP server

    “use the tools at sehn.ai”

    • Boundaries, shapes
    • Paths, connections
    • What’s in front
    • Rotate, count
  3. 03

    New answers

    About shape, position and connection, that it could not give before.

Illustrative examples

What we model

How the brain represents space, open to LLMs

Not just the first layers of the eye. We model a large part of the human visual system, grounded in decades of vision neuroscience, and make its picture of space readable for any LLM.

Latera visual-spatial model of its own.

First evidence

Five days in, agents trace paths they couldn’t before.

200 new mazes: which entrance reaches the exit? We generated them after building the tools. Nobody looked at them, nothing was tuned on them, and the analysis was written down before the run.

Claude Haiku 5.5 on 200 new mazes our tools never saw

  • Answers directly 63/200 · 32%
  • Agent, can zoom in 67/200 · 34%
  • Agent + sehn.ai tools 173/200 · 87%

Guessing scores 30%. On the 10 BabyVision puzzles the tools were built on: 10/12 with tools, 5/12 zooming, 4/30 directly.

Held up on its first real test94% on familiar maze styles, 79% on styles we never built on. Weak spot: colour-inverted regions (11/25). One task family, week one.

Where it matters

One way of seeing, many markets

The same sense of space helps wherever a machine has to understand a room, a drawing, a face or a screen.

Robotics

A free path through a cluttered room.

Engineering firms

Follow a pipe or wire through a drawing.

Image & video generation

Surfaces and depth that stay consistent.

Live avatars

A hand in front of a face, an object held up.

Live video

Follow what moves and connects across frames.

Working nowconnections and filtering, 87% on 200 unseen mazes
Built, in testingboundaries, surfaces, depth
Nextreal photos, drawings, live video

The name

Why we call it sehn

sehn means “to see” in Austrian German.

Diana is Austrian. For her, seeing is the most important part of a good life, and she loves that agents and LLMs can learn to see. Anton’s passion has always been vision: how we see, and how a machine could see the same way.

Both have loved Magic Eye pictures since they were children. A flat pattern, and suddenly depth appears. That moment is what we give to AI.

Autostereogram in plum, rose and cream dots; with relaxed eyes a hidden word floats out in 3D.

Try it: look through the picture, as if at something far behind it. One word floats out. When the two dots turn into three, you’re close.

How to see it

Bring your face close to the screen, so the picture is a blur. Then slowly lean back with relaxed eyes, as if you were looking at a distant point behind the screen. Give it 10 to 30 seconds.

Team

The two people behind sehn

A clinical psychologist who works with inner images, and an engineer who builds vision systems modelled on the brain.

Portrait of Anton Vattay

Anton Vattay

Co-founder · CTO · Research

  • 20 years of engineering
  • Vision AI at Toren AI and VEDO
  • Brain-inspired models of retina and cortex
  • Model training and deployment
Portrait of Dr. Diana Schaffer

Dr. Diana Schaffer

Co-founder · CEO

  • Clinical psychologist, PhD; studied computer science
  • 10+ years of practice in EMDR and trauma-reprocessing therapy
  • Organisational development, built high-performance teams
  • Former CEO of a medical center

How we got here

  1. 01In therapy, inner images change emotions; words alone rarely do.
  2. 02We tried to teach LLMs imagery therapy; they couldn’t understand an image.
  3. 03Anton saw the same blind spot in every vision model he built.
  4. 04sehn.ai: give AI the non-verbal reasoning it is missing.

Early access

Every AI that sees should see like we do.

sehn.ai is in early access. Tell us what your agent should see, and we’ll send you an API key to try it.

anton@mindsprings.ai diana@mindsprings.ai

Ask for an API key

Opens your email app with this text filled in. Nothing is sent until you press send.