· 4 min read
What AI Can Really Do in Task Management – and What It Can't
AI in to-do apps: where language models genuinely help with capturing, breaking down and prioritising tasks, where they fail (time estimates, hallucinations, priorities without context), and how to use them well.
"With AI" is now printed on every second productivity app. Mostly it means a chat window that summarises whatever you paste into it. That is nice, but it is not what makes task management hard. Three things are hard: capturing everything that is pending, finding the right order, and starting. Language models can really help with two of these, only indirectly with one, and they can do damage with all three if used badly.
What works well
Pulling tasks out of text. This is the strength that is advertised least and delivers most. Meeting minutes, a long email, a note from the fridge: language models are very good at reading the actual actions out of it, phrasing them as short sentences and recognising dates ("by Friday" becomes a concrete date). That does not just save typing. It lowers the barrier to capturing so far that things actually make it into the list instead of staying in your head.
Breaking tasks down. "Organise the move" is not a task, it is a project. A model can turn it into five concrete first steps in seconds. They are not always perfect, but they are a start, and the start is psychologically the most expensive part. If you can see the steps, you no longer have to work against the mountain.
Comparing, when there is context. This is the most interesting capability, and it has a condition. A model asked to rate a single task guesses. A model that sees thirty tasks at once, plus a paragraph about the person's situation ("freelance, two kids, money is tight right now, I can only concentrate in the evening"), can do something no rule can: recognise that the kids' swimming course secures evening childcare and is therefore not leisure but a precondition for working. That is not a magic trick, it is reasoning over connections you gave it.
Asking back. A good system recognises when it does not know what is meant. "Something with the car" has no priority until it is clear whether the garage or the insurance is meant. A model that asks a question at this point, instead of inventing a number, is worth more than one that always has an answer.
What does not work
Time estimates. Language models estimate effort about as badly as people do, and people systematically estimate too low. Daniel Kahneman and Amos Tversky described this in 1979 as the planning fallacy, and no model has overcome it. An AI estimate of "30 minutes" is an indication of the order of magnitude, not a basis for planning.
Priorities without context. Without knowing who the person is and what they want, a model rates by platitudes: taxes important, health important, tidying unimportant. That is not wrong, but it is not your priority. If you do not tell the AI what counts right now, you get the priorities of an average person, and there is no such person.
Invented details. Models like to fill in. A date that was not in the text, a person nobody mentioned. Every output that goes into a list should be seen first. Not out of distrust, but because a wrong due date in a list is more expensive than ten seconds of checking.
Starting. No model can make you start the tax return. It can lower the barrier: show the first step, put the task on top, provide the reason. But the decision stays with you, and a tool that promises otherwise is lying.
Two risks rarely talked about
Trust without checking. An order with numbers looks objective. It is not. It is the opinion of a model based on what it knows. A good tool therefore shows the reasoning for every rating so that you can disagree. A number without a reason is an oracle, and oracles make people passive.
Data. Tasks are personal. Anyone sending tasks and a paragraph about their life to a model should know where to. The questions are simple: which provider, is it used for training, can I switch it off. If you cannot get these questions answered, do not use the feature.
How to use AI here sensibly
Give it context once, in one paragraph: what you do, what is tight right now, what you want to reach, when you can work. Let it capture, break down and compare. Read the reasons, and disagree when it is wrong. Keep the decision. And do not expect any order to take the starting off your hands. It can only make it easier. Which, admittedly, is not nothing.
More from the blog
-
Shared Lists, Clearly Assigned Tasks: How Families and Small Teams Share the Mental Load
Mental load is not created by doing, but by keeping track. What research on invisible work and diffusion of responsibility shows, and how shared lists with clear ownership really lighten the load.
-
Procrastination Is Not a Character Flaw: What Psychology Knows About the Tasks We Put Off
Why we postpone unpleasant tasks, what research from Piers Steel to Timothy Pychyl says about it, and five interventions that demonstrably help – without more discipline.
-
Eisenhower, GTD, Eat the Frog: What the Classics Deliver – and Where They Break Down
The Eisenhower matrix, Getting Things Done, Eat the Frog, the Ivy Lee method: a sober comparison of the well-known task methods, their strengths, and the one gap they all share.