MonDive#45: ChatGPT 5.6: The Complete Guide

Learn how to choose the right GPT-5.6 model, and use Work, Sites, plugins, and skills effectively.

Smarter with AI banner

Welcome to the MonDive

Today in MonDive, we’re looking at GPT-5.6, OpenAI’s new three-model family built for different levels of work, from difficult reasoning to fast, high-volume tasks.

We’ll break down the differences between Sol, Terra, and Luna, compare their benchmarks and pricing, see how Sol stacks up against other frontier models, and explore the new ChatGPT features and real-world use cases.

Alright, let’s dive in.

Your chance to be a part of the FREE Claude AI Mastery Workshop that can 10x your Output

Claude is currently the most powerful tool of 2026. It's been launching new features every week- Skills, Connectors, Cowork, vibe coding. Yet almost no one knows how to actually use them.

Our expert mentors have condensed 800+ hours of Claude research, articles, YouTube content and real-world practice into a focused 16-hour curriculum. Join the 2-Day Claude AI Mastery Workshop: a live, end-to-end deep dive into Claude plus 10+ AI tools, LLMs and workflows.

It would be silly not to SIGN UP

You will learn how to:
- master Claude's three modes: Chat, Cowork and Code.
- Set up Skills, Connectors and Plug-ins to automate your desktop, Notion and files.
- Vibe code apps and dashboards without writing code & 10+ AI tools and workflows that pair with Claude.

🧠 Saturday & Sunday
🕜 10 AM – 7 PM EST

What GPT-5.6 Actually Is

GPT-5.6 is not one model with a few speed settings. It is a family of three models, each designed for a different balance of capability, speed, and cost.

The generation number tells you how new the family is. The name tells you where that model sits in the lineup:

  • Sol — The high-capability variant, flagship model for the most demanding work

  • Terra — The balanced variant, designed for everyday professional tasks

  • Luna —Luna — The efficient variant, built for fast, affordable, high-volume work

OpenAI says these names are durable capability tiers, so future generations can improve without changing what each tier represents. The shift is simple: instead of forcing one model to handle every kind of job, GPT-5.6 gives each level of work its own model.

Choosing Between Sol, Terra, and Luna

The difference becomes clearer once you match each model to the kind of work in front of you.

Sol: The High-Capability Tier

Sol is the premium choice when the brief is messy, the stakes are high, or the task needs deeper judgment before anything gets built.

Best for

  • Complex coding, debugging, and architecture decisions

  • Research that spans multiple files, sources, or tools

  • High-value work where accuracy and polish matter

Choose Sol if

  • The task is difficult, open-ended, or has many connected parts

  • You care more about the strongest result than the quickest one

Terra: The Balanced Tier

Terra is designed for serious daily work where you still need strong reasoning, but bringing in the flagship model would be unnecessary.

Best for

  • Everyday coding, debugging, and feature work

  • Research, writing, editing, and summarization

  • Documents, presentations, and business workflows

Choose Terra if

  • You want a capable default for regular professional work

  • The goal is clear, but the task still needs thoughtful execution

Luna: The Efficient Tier

Luna fits work that is already well defined and needs to be completed quickly across a large number of requests.

Best for

  • Sorting, tagging, and classifying information

  • Extracting or reformatting structured data

  • Processing batches of transcripts, messages, or documents

Choose Luna if

  • You need many similar tasks completed at scale

  • Speed and efficiency matter more than deep reasoning

Choosing the Right Effort for Each Model

The effort setting controls how much time and computation the selected model spends on a task. In ChatGPT Work and Codex, the main options are Light/Low, Medium, High, Extra High, and Max. Higher settings can improve planning and checking, but they also take longer and use more tokens. Ultra is different: it can split a task across multiple agents instead of simply making one agent think longer.

Model

Recommended starting point

When to increase it

Sol

Medium

Move to High or Extra High for difficult work. Use Max only when quality matters more than time or usage.

Terra

Medium

High or Extra High can work well when the task is clear but needs more planning, accuracy, or verification.

Luna

Light/Low or Medium

Raise it for structured tasks that need extra care. If you regularly need Extra High or Max, moving to Terra or Sol may make more sense.

Using Terra at High or Extra High can sometimes be a smarter choice than using Sol at Low, especially for a well-defined task where you want more reasoning without moving to the flagship model. But there is no fixed rule that one will always outperform the other, so OpenAI recommends testing familiar tasks and using the lowest effort that produces the result you need.

Max is available to GPT-5.6 users in Work and Codex once enabled in settings. In ChatGPT Work, Ultra is limited to Pro and Enterprise, so it is not included with Plus; in Codex, Ultra is available on Plus and higher plans. Most tasks will not need either setting.

How Sol, Terra, and Luna Compare on Benchmarks

OpenAI’s published results show that the three models stay fairly close on some professional and coding tasks. The differences become much larger when the work requires stronger computer control or reliable reasoning across very large amounts of information.

Benchmark

GPT-5.6 Sol

GPT-5.6 Terra

GPT-5.6 Luna

What it tests

Agents’ Last Exam

52.7%

50.4%

50.3%

Long professional workflows

Coding Agent Index v1.1

80

77.4

74.6

Coding across implementation, terminals, and real codebases

OSWorld 2.0

62.6%

50.2%

45.6%

Using and navigating real computer environments

MRCR v2, 256K–512K

91.5%

89.6%

41.3%

Retrieving the right information from very large contexts

  • Terra stays surprisingly close to Sol on professional workflows, coding, and long-context retrieval.

  • Luna remains competitive on structured agent and coding tasks, but falls further behind when computer use or difficult long-context work is involved.

  • Sol has the clearest advantage in demanding tasks that require the model to navigate software, retain a lot of context, and make decisions across several steps.

These are OpenAI’s own published evaluations, so they are best read as a useful guide rather than a final verdict. Real-world performance will still depend on the task, prompt, effort setting, and tools involved.

How Much Sol, Terra, and Luna Cost

OpenAI’s pricing follows the model ladder closely: Terra costs half as much as Sol, while Luna costs one-fifth as much. The rates below are the standard API prices per one million text tokens for requests with up to 272,000 input tokens.

Model

Input

Cached input

Cache write

Output

GPT-5.6 Sol

$5.00

$0.50

$6.25

$30.00

GPT-5.6 Terra

$2.50

$0.25

$3.125

$15.00

GPT-5.6 Luna

$1.00

$0.10

$1.25

$6.00

  • Terra is priced at exactly half of Sol, making it easier to use regularly without dropping to the lowest-capability tier.

  • Luna is the cheapest option, with standard input and output rates that are 80% lower than Sol.

  • For prompts above 272,000 input tokens, the full request is charged at twice the normal input rate and 1.5 times the normal output rate.

These figures cover model tokens only. Features such as web search, file search, and hosted computing can add separate tool charges to the final cost.

How Sol Stacks Up Against Fable, Grok, and Muse

Sol sits near the top of the current frontier, but it does not win every comparison. Artificial Analysis places Claude Fable 5 slightly ahead on broad intelligence, while Sol leads its coding-agent evaluation and completes the wider test suite at a much lower estimated cost. Grok 4.5 and Meta’s current model, Muse Spark 1.1, sit further down the overall ranking but remain competitive because they deliver strong results with far less spending.

Model

Intelligence Index

Estimated cost per task

Where it stands

GPT-5.6 Sol (Max)

59

$1.04

Nearly matches Fable and leads the Coding Agent Index

Claude Fable 5 (Max)

60

$2.75

Highest broad intelligence score and stronger on some deep engineering tasks

Grok 4.5 (High)

54

$0.31

Strong coding and terminal performance at a much lower cost

Muse Spark 1.1 (XHigh)

51

$0.26

Fast, affordable, and built for multimodal and agentic work

  • Against Fable 5, Sol delivers almost the same broad intelligence at roughly one-third of the estimated task cost, although Fable still holds a clear advantage on some repo-level coding and analytical-work evaluations.

  • Against Grok 4.5, Sol has the higher overall and coding-agent scores, while Grok stands out as a cheaper, highly efficient option for coding-heavy and terminal-based workflows.

  • Against Muse Spark 1.1, Sol remains the more capable general frontier model. Muse’s strength is different: Meta built it as a fast multimodal and agentic model for Meta AI and its developer API.

These are results from a shared independent evaluation, so they offer a useful snapshot rather than a final answer for every workflow. The best model can still change depending on the tools, reasoning setting, and type of task involved.

New ChatGPT Features

GPT-5.6 is only part of the update. ChatGPT now has new tools for completing full projects, publishing web apps, automating recurring work, and connecting the services you already use.

ChatGPT Work

Chat is for questions, while Codex remains focused on software development. Work is for broader professional tasks where you want to hand off a goal and receive a finished result.

You can connect tools such as your calendar, Slack, and Google Drive, add relevant files or project context, and let ChatGPT gather the information, create a plan, and complete the work. It can produce documents, presentations, spreadsheets, reports, dashboards, and other deliverables while letting you follow its progress, answer questions, and approve important actions.

People can use Work to:

  • Prepare a daily briefing from meetings, messages, and files

  • Turn research or brand material into a campaign plan and presentation

  • Combine product metrics, customer feedback, and team discussions into one report

ChatGPT Sites

Sites lets you turn a description into a working website or lightweight web app without setting up a separate hosting service.

ChatGPT creates the project, writes the code, and gives you a private preview to test. You can then request changes in plain English, leave targeted comments inside the preview, and publish the final result through a shareable link.

Sites can be useful for:

  • Landing pages and product launches

  • Dashboards and interactive reports

  • Project trackers and internal portals

  • Prototypes, small web tools, and browser games

The result is not limited to code inside a conversation. It becomes a hosted Site that other people can open and use.

Scheduled Tasks

Scheduled Tasks removes the need to repeat the same prompt every day or week.

You can ask ChatGPT to run something once, repeat it on a schedule, or monitor for changes and notify you only when something important happens. A dedicated Scheduled page lets you review upcoming runs and pause, edit, resume, or delete existing tasks.

People can use it to:

  • Receive a briefing before the workday begins

  • Summarize weekly feedback from Slack, Notion, or Drive

  • Prepare recurring reports and post them to the right team channel

  • Monitor information and alert you when there is a meaningful update

For Business and Enterprise teams, a repeatable workflow can also be turned into a Workspace Agent that runs in the cloud and can be shared through ChatGPT or Slack.

Plugins and Skills

Plugins expand what ChatGPT can access and do. They can connect external services, provide specialized tools, or bundle several capabilities into one reusable workflow. Plugins are available in Work and, on supported surfaces, Codex.

Plugin

What people can use it for

Google Drive

Find and work across Drive, Docs, Sheets, and Slides

Gmail

Search messages, organize email, draft replies, and identify follow-ups

Slack

Summarize channels, recover project context, and draft responses

GitHub

Work with repositories, issues, pull requests, and development activity

Codex Security

Scan authorized code, validate findings, and prepare reviewed fixes

A plugin gives ChatGPT access to tools and information.

A Skill gives it a repeatable way of working.

Skills can store instructions, examples, reference files, and scripts for a specific workflow. You could create one for your newsletter voice, presentation style, research process, reporting format, or publishing checklist, so ChatGPT follows the same method without needing the full instructions pasted into every conversation.

GPT-5.6 Use Cases

The most interesting part of GPT-5.6 is the range of work people are already pointing it at. The examples move from difficult research and business operations to physical automation, interactive learning, and full 3D applications.

  • Scientific research — A mathematician used GPT‑5.6 in Codex to explore a problem in algebraic geometry that had resisted previous methods for three years. It found a new direction that helped the team disprove the conjecture they had been working on.

  • Small-business operations — A cereal company turned historical launches, retailer deadlines, contacts, and brand guidelines into a custom command center for managing seasonal product launches.

  • Farming and physical automation — A farmer used Codex to map field boundaries and automate greenhouse ventilation with electric motors, despite not being a software engineer.

  • Complex maps and simulations — Developers created a Google Earth-style globe, procedurally growing cities, and a detailed voxel version of Manhattan that ran as a long-term autonomous build.

  • Interactive learning tools — GPT-5.6 generated visual explainers such as wave simulations, spirographs, and tokenizers, letting people learn by changing inputs and watching the result update.

  • Polished apps and games — It can build browser games and front-end experiences, inspect the rendered output, and improve the visuals and interactions instead of stopping as soon as the code runs.

How did you feel about today’s MonDive?

Was this guide easy to follow?

Login or Subscribe to participate in polls.

Know someone who may be interested?

And that's a wrap on today's MonDive!

Reply

or to participate.