CASE 01EvaluationSol and Luna in a landing-page prompt comparison
A creator added both GPT-6 models to a same-prompt landing-page gallery; the published results are a comparison, not an independently reproduced build.
RECENTLY UPDATED / (SHANGHAI TIME)
LIN YUEJI / ARCHIVE
Explore public GPT-6 Sol and Luna cases across web builds, games, agents, video workflows, and evaluations, with model roles and reported limits made clear.
USE CASE MAP / 01
Only public tasks, evaluations, and failures clearly attributed to GPT-6 Sol or Luna are included. Each case links to its creator and source; most results are author or platform reports, not independent reproductions.
HOW TO USE / 02
Check the model variant and task on each card, then open the source for implementation details, tool dependencies, and evidence limits.
Start with coding, agents, creative work, evaluation, or another task family.
Use titles, methods, or creators to find the closest case to your current goal.
Select the creator to review the full demo, limits, context, and original explanation.
Choose a method worth testing, connect to the model, and turn it into your workflow.
REAL CREATOR CASES / 03
Up to four cases per row. Select a creator to open the original case; images and videos load only when needed.
58 cases
CASE 01EvaluationA creator added both GPT-6 models to a same-prompt landing-page gallery; the published results are a comparison, not an independently reproduced build.
The creator compared Sol with the earlier Opus 5 using the same prompt and max effort; this is not an Opus 5.5 comparison.
An API platform showed a 3D quality and cost comparison; the result is platform-reported and was not reproduced.
CASE 04EvaluationAn API platform gave Sol and Opus 5.5 the same brief for a looping 3D pelican animation; the aesthetic verdict is the author’s.
CASE 05EvaluationA same-task visual comparison pits GPT-6 Luna Max against GPT-5.6 Luna Max; source media shows the output, without code verification.
The post compares a specific SVG task under Sol xHigh and Astra High; visual output is shown, while source code remains unverified.
CASE 07LimitThe creator reports that Sol produced a static frame without the requested render loop or camera motion; this is a reported incomplete result.
CASE 08EvaluationHiggsfield compared Sol and Opus 5.5 in an Unreal Engine game build; the project and playability were not independently tested.
The creator shows a short Luna Max game demonstration; the game was not independently run.
The creator demonstrates a Minecraft-style build attributed to Sol; source code and playability were not verified.
JetBrains showcased a Sol-assisted Neon Exit game made by an employee; this is a platform-published demonstration.
CASE 12DemoA creator says Luna made a model introduction video; the production chain is undisclosed and does not establish native video generation.
CASE 13EvaluationMirage showed Tesseract-assisted results for Sol and Astra; the rendering role belongs to the video tool, and the comparison was not reproduced.
A creator demonstrated a few-prompt workflow combining Sol with HyperFrames; the post does not show native Sol video generation.
CASE 15DemoTesseract’s co-founder shows Sol-assisted brand and aspect-ratio edits; the result is a product demonstration, not independent validation.
CASE 16BenchmarkBrowser Use published a benchmark of long, difficult browser-agent tasks; its cost and performance claims are platform-reported.
The author reports switching existing scheduled tasks to Sol; replacing another API with Luna remained a future plan.
CASE 18EvaluationA researcher reports improved Sol tool use and weaker Luna performance relative to the prior generation; the evaluation was not reproduced.
CASE 19BenchmarkThe Next.js team reports a 97% success rate for Sol on its framework eval; this does not measure all software development tasks.
CASE 20EvaluationA developer reports both suites passing with Luna at $0.03 versus $0.15 for the prior model; these are author-recorded costs.
CASE 21DemoA creator instructed Sol to open Paint in the browser and draw a portrait; the post shows the reported result.
A publisher asked models for 50 scripts across ten genres and used three AI judges; Sol’s reported result reflects this test, not audience response.
CASE 23EvaluationPerplexity reports a WANDR research evaluation and plans Sol for its Light orchestrator; planned routing is separate from measured results.
A researcher reports Luna failing to open a stove and place a moka pot; the task name comes from a quoted Astra experiment.
CASE 25EvaluationBridgeBench compared four models on one lava-lamp prompt; the posted Sol/Luna times and costs are test-specific.
CASE 26BenchmarkThe publisher reports Sol Max finding 29.3 of 105 seeded bugs at a lower estimated cost than several peers; quality fell in this test.
The author says Luna ignored a GitHub integration and opened a browser instead during a repository review, across four harnesses.
CASE 28DemoHiggsfield showed a Sol versus Opus 5.5 dogsled game comparison; the platform supplied the creation environment.
CASE 29EvaluationBridgeBench reports Sol completing the shared scene prompt in 44 seconds at $0.07; this is one visual comparison.
CASE 30DemoHiggsfield compared Sol and Opus 5.5 on a battlefield-scale game prototype; the result is a platform demo.
CASE 31DemoA creator reports using Sol with Combos CLI for a 2–4 player board game; the external tool supplied assets, networking, and publishing.
CASE 32BenchmarkRoboflow reports lower extraction, counting, and reasoning scores than GPT-5.6 Luna, while detection improves slightly.
CASE 33LimitIn a peacock animation test, the creator says Luna inserted an image and spinning wheel circles instead of constructing the requested animation.
CASE 34EvaluationThe prompt asked for a monster that follows the cursor and reacts to feeding; the publisher compared both GPT-6 variants and reported Luna finishing first.
CASE 35EvaluationA same-prompt test compared playable pinball builds; the publisher reports Luna at 6:22 and Sol at 17:26.
CASE 36DemoHiggsfield reports a single HTML project with browser-rendered frames and timeline controls; Sol’s role was code generation.
CASE 37EvaluationThe creator reports a three-planet site from Sol in ten minutes; the comparison used a different subscription plan for Opus.
CASE 38EvaluationA creator compared four models on a Hermes Agent kanban app prompt and ranked Sol last for frontend design.
CASE 39EvaluationOpenDesign compared Sol and Opus 5.5 on the same city prompt and reported distinct visual styles.
CASE 40BenchmarkA 32-CVE benchmark reports 68.8% recall for Sol and 53.1% for Luna; both trail their GPT-5.6 predecessors on recall.
CASE 41LimitA same-prompt Blender test reports faster, cheaper Sol output with visible geometry and rigging problems.
CASE 42BenchmarkGertLabs reports Sol near Opus 5.5 within its margin of error and Luna slightly ahead of 5.6 Luna in its custom coding evaluation.
CASE 43DemoThe creator used Astra for planning and Sol for coding, with Combos CLI and other models for 3D assets, sound, multiplayer, and publishing.
CASE 44DemoA creator built a co-watch site using ChatGPT Voice and WebMCP; Sol comments on a synced video during playback.
CASE 45EvaluationThe same prompt asked four models to write a product update with a UI screenshot; the author estimated Sol’s cost at $3.10 and preferred another result.
CASE 46BenchmarkMazeBench reports Sol scoring 1% on its difficult spatial tasks; this is a narrow benchmark result.
CASE 47BenchmarkArena reports Sol Max at fourth place in its WebDev ranking; the score reflects that evaluation’s prompts and voters.
CASE 48BenchmarkA Rails agent evaluation reports Luna Max completing 18% of tasks for $11 and compares it with Sol in the same chart.
Showing 48 / 58
BUILD WITH EVOLINK / 04
Check the current model catalog and integration options on EvoLink, then test a relevant case on your own task.