2026-03-18

Space Bunny Alpha Benchmark Notes: Coding & Long Context

Independent notes on Space Bunny Alpha preview behavior for coding tasks and long-context reading.

Independent benchmark notes

These notes are informal observations from the Space Bunny Alpha demo playground — not vendor claims. Preview models shift; re-run your own evals before production decisions.

Coding tasks

On medium TypeScript refactors and React debugging, the model generally follows instructions when prompts specify output format. It benefits from temperature around 0.2–0.5 for precise edits and higher creativity for brainstorming architecture options.

Long context reading

With large pasted corpora, ask for citations back to section headers or function names. That validates whether the model is using the provided context rather than improvising. Track credit burn: long prompts dominate cost even when answers are short.

Multimodal

Screenshot-to-code quality depends on image clarity. Prefer crisp UI captures under 4MB. Combine the image with a short constraint list (stack, spacing, accessibility) for better playground results.

Ready to try it? Open the playground or get 100 free credits.