A research lab for long-horizon AI

Intelligence over
longer horizons.

We build interactive environments for agents that
reason, act, and adapt. Our focus is the work that
unfolds over many steps.

Our research direction
State / action / consequence
General Context Labs, Inc.Beyond a single interaction.

The next frontier is
following through.

Many useful problems cannot be solved in a single response. They require an agent to make a plan, act on an environment, interpret feedback, and revise its approach—while keeping a larger objective in view.

We study the environments and evaluations that make this kind of capability possible to develop and measure. Our central question: how do we build agents that remain effective as tasks get longer, state gets richer, and decisions compound?

Long-horizon reasoning

Planning across many dependent steps. Maintaining context, recovering from mistakes, and carrying work through to completion.

Interactive environments

Worlds with persistent state and meaningful consequences, where agents must observe, act, and adapt as a task unfolds.

Grounded evaluation

Measuring what an agent accomplishes. Testing behavior in the environment, beyond the plausibility of its final response.

Research & writing.

Benchmark

Three.js Asset Generation Benchmark

A benchmark for structurally correct, visually convincing 3D assets generated by frontier models. We examine whether generated objects satisfy their prompts through executable checks and geometric verification.

Hard tasks.
Observable outcomes.

Our research takes shape in engineering tasks, spatial evaluations, and generation trajectories. Each offers a different view of how agents reason, build, and improve.

01

Long-horizon tasks

Substantial engineering problems that demand sustained reasoning, implementation, testing, and refinement. Tasks span compilers, systems programming, scientific computing, machine learning, and visual reconstruction.

  • Compile Python to WebAssembly
  • Build an H.264 decoder in Zig or a SAT solver with DRAT proofs
  • Optimize BLAS, FFT, compression, and scientific libraries
  • Train a CIFAR-10 classifier using local learning
  • Reconstruct a 3D street or build a cycle-faithful emulator

Evaluated through task-specific tests, reference behavior, held-out data, and correctness-gated performance.

02

3D evaluations

Three.js Asset Generation (3JAG) evaluates whether generated objects execute, render correctly, and satisfy the spatial relationships in a brief.

  • Executable geometry and rendering checks
  • Contact, clearance, connectedness, and grounding
  • Task-specific geometric constraints

Deterministic measurements connect each verdict to the generated artifact.

Read the benchmark research
03

Game generation trajectories

The process behind a working game: creator requests, iterative code changes, and playable versions from the first build to the final result.

  • Sequences of generation and editing requests
  • Before-and-after builds and source changes
  • Playable outcomes across successive revisions

A view into how agents respond to feedback and develop an artifact over time.

Games were the beginning.

A broader horizon

Instaplay turns ideas into playable worlds. Games brought code, visual reasoning, and dynamic environments together—and gave our research a starting point.

Instaplay Ideas become playable worlds. Explore Instaplay (opens in a new tab)

Longer horizons.
Broader possibilities.

Research at the intersection of agents, environments, and evaluation.

Discuss research & data