Mood tracking apps produce charts that look like measurement, and people draw conclusions from their shape. Several properties of self-recorded data make those conclusions unreliable in specific and predictable ways.
Entries are not evenly distributed
People log more often when something notable happens and less often during unremarkable stretches.
The record therefore over-represents peaks and troughs relative to the ordinary middle, which is exactly the region that dominates most weeks.
A chart built from that record makes mood look more volatile than it is, because the flat periods are thinly sampled rather than absent.
The act of rating changes the rating
Assigning a number to a mood requires attending to it, and attention to mood alters mood, particularly for low states.
Some people find this helpful, since naming a state reduces its intensity; others find that repeated checking amplifies whatever is being checked.
Which effect dominates varies between individuals, so the tool is not neutral with respect to what it measures. A month of logging is itself an intervention, and its record cannot be read as though it were passive observation.
Scales drift over time
A rating of four means whatever the person's recent range has made it mean, and that reference point shifts as circumstances change.
After a difficult period, an average day may be rated highly; after a good stretch, the same day may be rated lower.
This drift makes long-run comparisons across months less trustworthy than comparisons within a few weeks, which is the opposite of how the charts are usually read.
Correlations found in the data are weak evidence
Apps often surface associations between logged factors and mood, such as sleep, exercise or caffeine.
With many factors and many days, some associations appear by chance alone, and the app has no way to distinguish those from real ones.
Direction is also ambiguous: low mood reduces exercise as readily as exercise raises mood, and a record of what happened on the same day cannot separate the two.
Where the record does add value
Tracking is most useful as a memory aid for a clinical conversation, since recall of the past few weeks is unreliable and shaped by current state.
A record of when difficulty began, what changed around it, and how long episodes lasted is genuinely informative to a clinician.
Used that way it supports assessment rather than substituting for it, and a sustained low period recorded in an app is a reason to seek help rather than to keep logging.