Problem guide
How to Create an Accurate Interactive Sankey Diagram
Use a Sankey only when you really need to show flow from one stage to another. Give every node a stable ID, make sure totals balance, roll tiny paths into honest “Other” groups instead of deleting them, and let people click a path to follow it.
Written for: Analysts, researchers, operations teams, and data storytellers
Use a Sankey only when you really need to show flow from one stage to another. Give every node a stable ID, make sure totals balance, roll tiny paths into honest “Other” groups instead of deleting them, and let people click a path to follow it.
Use a Sankey only when the flow is the message
A Sankey diagram is appropriate when the audience needs to understand how a quantity moves through stages, splits among destinations, or combines from sources. It is not simply a decorative alternative to a funnel. The width of every link should represent a comparable amount, and the stages should form a meaningful sequence.
Good questions include:
- Where does funding originate, which programs receive it, and what outcomes does it support?
- How do leads move from acquisition channels through qualification to final status?
- Which origins feed major entry points and final destinations?
- How does energy enter, transform, and leave a system?
A Sankey is a poor choice when the data is merely a set of unrelated category pairs, when exact ranking matters more than pathways, or when the audience must compare dozens of tiny flows. In those cases, a matrix, stacked bar, alluvial small multiple, or table may be clearer.
Define stages and node identity explicitly
Write the stage model before preparing links. Each node needs a stable ID, a display label, and a stage. Do not rely on a repeated label such as “Other” to identify a node, because “Other origins” and “Other destinations” are different entities.
A reliable node table might contain:
| node_id | label | stage | order |
|---|---|---|---|
| origin_web | Website content | 1 | 1 |
| status_web_qualified | Qualified | 2 | 1 |
| outcome_won | Won | 3 | 1 |
The link table should use IDs:
| source_id | target_id | value | flow_type |
|---|---|---|---|
| origin_web | status_web_qualified | 220 | observed |
Separate IDs from labels so nodes can share reader-facing wording without merging accidentally. Store values as nonnegative numbers unless you have designed a specialized signed-flow treatment.
Make sure the flows add up—or explain why they do not
For an ordinary Sankey, what flows into an intermediate node should match what flows out. If 220 units enter “Qualified,” the outgoing paths should account for those 220 units. If some are still active, lost, or unresolved, show that instead of making them disappear.
Calculate these checks outside the visualization:
- sum of outgoing links by source;
- sum of incoming links by target;
- difference at every intermediate node;
- total entering and leaving each stage;
- count and value of omitted links.
Sometimes a node represents a stock, a loss, or an incomplete observation window, so exact conservation is not expected. In that case, create an explicit “Unresolved,” “Still active,” “Loss,” or “Not observed” destination. Invisible disappearance makes readers assume a data error.
Simplify the diagram without changing the totals
Real flow data often contains hundreds of small paths. Filtering them away can make the diagram legible but can also shrink node totals and distort the story. Use a transparent reduction strategy.
One strong approach is:
- retain all major nodes required by the narrative;
- keep flows above a documented threshold;
- aggregate excluded nodes into stage-specific “Other” nodes;
- add remainder links so displayed node totals match the full data;
- style remainder flows neutrally and explain the rule.
Do not use one universal “Other” node across multiple stages. A flow from “Other origins” to “Other destinations” may be necessary, but it should not create an artificial pathway that did not exist in the source. Aggregate at the link level with care.
If the full network remains too dense, offer a high-level default and let readers select one source or node to isolate its pathways. A readable truthful subset is better than an unreadable hairball.
Choose node order deliberately
Automatic layout algorithms minimize link crossings, but their order can shift when data changes. That makes recurring reports difficult to compare. Provide a stable order based on stage logic, business sequence, geography, or total volume.
Keep related nodes adjacent. When categories have an inherent order—awareness, consideration, decision; low, medium, high—preserve it. Do not sort each stage independently if that breaks visual continuity.
Document whether node height represents total incoming, outgoing, or the maximum of both. Most Sankey implementations calculate node size from links, so your reconciliation rules determine what readers see.
Use color to identify meaning, not every link
Color can encode source category, destination outcome, status, or flow type. Pick one semantic rule and maintain it. Coloring every link uniquely creates a legend problem and makes overlaps harder to interpret.
Common strategies include:
- links inherit the source color, useful when tracking origins;
- links inherit the target color, useful when outcomes matter;
- nodes use category colors while links remain translucent gray;
- highlighted selections use color and all other links recede.
Ensure adjacent colors remain distinguishable for common forms of color-vision deficiency, and do not depend on red versus green alone. Use labels and selection states to reinforce meaning.
Use interaction to make a dense flow easier to follow
A Sankey benefits from interaction when it helps the reader trace a path. Useful behavior includes:
- hover, focus, or click highlighting of all links connected to a node;
- a persistent detail panel with incoming, outgoing, and share values;
- selection of one origin, destination, or category;
- switching between amount and percentage while clearly changing labels;
- a searchable node list for large diagrams;
- a table of flows for exact values.
Click should lock a selection so touch and keyboard users can inspect it. Provide a visible reset. Tooltips should state the complete path or at least source, target, amount, unit, and percentage denominator.
Avoid animated particles unless they add real explanatory value. Movement can imply speed or time even when link width represents only volume.
Design labels and mobile behavior
Long node labels are often the limiting factor. Reserve space outside each stage, wrap labels predictably, and show totals with the label when scale matters. Abbreviations should have an accessible full form.
On a phone, a three-stage horizontal diagram may become too compressed. Options include horizontal scrolling with clear cues, a vertically oriented flow, stage-by-stage drilldown, or a simplified selection view. Do not simply shrink all labels and links.
Test a node with one link, a node with many links, the largest value, the smallest retained value, and a stage with many “Other” records. Ensure tooltips remain within the viewport.
Build the data preparation separately from the diagram
Prepare a canonical link table in SQL or Python. Apply stage mapping, deduplication, aggregation, thresholds, and remainder logic there. Keep the visualization code focused on layout and interaction.
In Rhubarb, you can start by asking the assistant to find and visualize the flow you care about—“Show how leads move from acquisition channel to final status” or “Trace funding from source to program to outcome.” It can draft the query that produces source_id, target_id, value, stage metadata, labels, and optional detail fields, then build the Sankey. You should still inspect conservation, aggregation, and node ordering before treating the result as final.
Save a version before changing aggregation rules. A visual refinement and a methodology change should not be merged into one opaque edit.
Check that the final flows add up
Before publication, reconcile every stage total with an independent query. Confirm that aggregated “Other” paths preserve the full amount, that no source-target pair is duplicated unintentionally, and that filters do not remove a link while retaining an inconsistent node total.
Ask a new reader to trace one major path and explain what link width means. If they interpret width as probability, time, or count when it actually represents dollars, strengthen the title, unit, and tooltip.
A successful Sankey does three things at once: it keeps the system’s accounting honest, makes the dominant pathways visible, and lets the reader investigate without being buried in every minor connection.
Frequently asked questions
Do Sankey flows have to balance?
Intermediate nodes should normally reconcile incoming and outgoing values. When stock, loss, unresolved cases, or an incomplete period prevents balance, represent and explain that remainder explicitly.
How should small Sankey flows be removed?
Aggregate excluded nodes into stage-specific Other nodes and preserve the omitted amount through remainder links. Do not simply delete small flows and leave misleading node totals.
Can a Sankey show negative values?
Ordinary Sankey layouts assume nonnegative link widths. Signed flows require a specially designed encoding and should not be passed into a standard Sankey implementation.