Essay

Thinking With LLMs - How To Use Them To Think Better, Not Less

Most people use LLMs to skip thinking. A few use them to think harder than they could alone. The difference is not the tool, the model, or the prompt - it is whether you are outsourcing cognition or amplifying it. Here i

Most people use LLMs to skip thinking. A few use them to think harder than they could alone. The difference is not the tool, the model, or the prompt - it is whether you are outsourcing cognition or amplifying it. Here is what the second mode actually looks like in practice.

Two people use the same LLM to work through the same problem. One finishes in twenty minutes with an answer they do not fully understand and could not defend if pressed. The other finishes in two hours with a position they have stress-tested from four angles and could explain to a hostile audience. Same tool, opposite outcomes.

The difference is not prompt engineering. It is not model choice. It is not which system they paid for. It is whether they were using the LLM to think or to avoid thinking. That distinction is the single most important variable in how much value you actually extract from these systems, and almost nobody in the productivity discourse talks about it.

There are roughly two ways to use an LLM for any cognitive task. The first mode treats it as an answer machine. You have a question, you type the question, you take the answer, you move on. The output looks like work. You feel productive. You ship something.

The second mode treats it as a thinking surface. You have a vague intuition, you externalize it, you let the model push back, you refine, you push back on the refinement, and somewhere inside that loop you actually figure out what you think. The output of mode two also looks like work, but it carries something the first mode never produces: a defensible position you actually own.

The first mode is what most people do most of the time. It is also what every “ten ChatGPT tricks” tutorial optimizes for. There is nothing wrong with mode one for trivial tasks where speed is the entire point and the cost of being wrong is near zero. Drafting a polite decline email. Reformatting a CSV. Boilerplate that someone has to write but nobody has to think about.

The problem starts when mode one quietly leaks into work that should have been mode two. Strategic decisions, research syntheses, performance reviews, technical tradeoff analyses, anything where being wrong costs something real. These all feel faster in mode one because they are faster. They are also worse. The speed is paid for in conclusions you cannot defend and reasoning you cannot reconstruct.

Vague prompts produce vague outputs. Everyone repeats this. Almost nobody acknowledges what it implies: the ceiling on what you can extract from an LLM is set by the quality of your thinking before you typed anything.