Dataclock API
The main chart function is dataclock().
Dataclock
- dataclocklib.charts.dataclock(data, date_column, agg_column=None, agg='count', mode='DAY_HOUR', cmap_name='RdYlGn_r', cmap_reverse=False, spine_color='darkslategrey', grid_color='darkslategrey', default_text=True, *, chart_title=None, chart_subtitle=None, chart_period=None, chart_source=None, **fig_kw)[source]
Create a data clock chart from a pandas DataFrame.
Data clocks visually summarise temporal data in two dimensions, revealing seasonal or cyclical patterns and trends over time. A data clock is a circular chart that divides a larger unit of time into rings and subdivides it by a smaller unit of time into wedges, creating a set of temporal bins.
TIP: Palettes - https://python-graph-gallery.com/color-palette-finder/
- Parameters:
data (DataFrame) – DataFrame containing data to visualise.
date_column (str) – Name of DataFrame naive datetime64 column, of any resolution (‘ns’, ‘us’, ‘ms’ or ‘s’).
agg_column (str, optional) – DataFrame Column to aggregate.
agg (Aggregation, optional) – Aggregation function; ‘count’, ‘max’, ‘mean’, ‘median’, ‘min’ & ‘sum’.
mode (Mode, optional) – A mode key representing the temporal bins used in the chart; ‘YEAR_MONTH’, ‘YEAR_WEEK’, ‘WEEK_DAY’, ‘DOW_HOUR’ & ‘DAY_HOUR’.
cmap_name (str, optional) – Name of a matplotlib/PyPalettes colormap, to symbolise the temporal bins; ‘RdYlGn_r’, ‘CMRmap_r’, ‘inferno_r’, ‘Alkalay2’, ‘viridis’, ‘a_palette’ etc.
cmap_reverse (bool, optional) – Reverse cmap colors flag.
spine_color (str, optional) – Name of color to style the polar axis spines.
grid_color (str, optional) – Name of color to style the polar axis grid lines.
default_text (bool, optional) – Flag to generating default chart annotations for the chart_title (‘Data Clock Chart’) and chart_subtitle (‘[agg] by [period] (rings) & [period] (wedges)’).
chart_title (str, optional) – Chart title.
chart_subtitle (str, optional) – Chart subtitle.
chart_period (str, optional) – Chart reporting period.
chart_source (str, optional) – Chart data source.
**fig_kw (Any) – Chart figure kwargs passed to pyplot.subplots; ‘figsize’ & ‘constrained_layout’ are always overridden, while ‘dpi’ (default 100) & any other kwargs are passed through.
- Raises:
AggregationColumnError – Missing agg_column for a non-count aggregation, a non-numeric agg_column for a non-count aggregation, or an agg_column named ‘ring’ or ‘wedge’.
AggregationFunctionError – Unexpected aggregation function value.
EmptyDataFrameError – Unexpected empty DataFrame.
KeyError – date_column or agg_column not in DataFrame.
MissingDatetimeError – Unexpected data[date_column] dtype, or data[date_column] contains NaT values.
ModeError – Unexpected mode value is passed.
- Returns:
A tuple containing a DataFrame with the aggregate values used to create the chart, the matplotlib chart Figure and Axes objects.
- Return type:
tuple[DataFrame, Figure, Axes]
- dataclocklib.charts.line_chart(data, date_column, agg_column=None, agg='count', mode='DAY_HOUR', default_text=True, *, chart_title=None, chart_subtitle=None, chart_period=None, chart_source=None, **fig_kw)[source]
Create a temporal line chart from a pandas DataFrame.
This function will divide a larger unit of time into rings and subdivide them by a smaller unit of time into wedges, creating temporal bins. The ring values will be represented as individual lines, with the aggregation values on the y-axis and wedges as the x-axis.
NOTE: fig_kw is accepted but currently unused; the figure is always created with figsize=(13.33, 7.5) & dpi=96.
- Parameters:
data (DataFrame) – DataFrame containing data to visualise.
date_column (str) – Name of DataFrame naive datetime64 column, of any resolution (‘ns’, ‘us’, ‘ms’ or ‘s’).
agg_column (str, optional) – DataFrame Column to aggregate.
agg (Aggregation, optional) – Aggregation function; ‘count’, ‘max’, ‘mean’, ‘median’, ‘min’ & ‘sum’.
mode (Mode, optional) – A mode key representing the temporal bins used in the chart; ‘YEAR_MONTH’, ‘YEAR_WEEK’, ‘WEEK_DAY’, ‘DOW_HOUR’ & ‘DAY_HOUR’.
default_text (bool, optional) – Flag to generating default chart annotations for the chart_title (‘Line Chart’) and chart_subtitle (‘[agg] by [period] & [period]’).
chart_title (str, optional) – Chart title.
chart_subtitle (str, optional) – Chart subtitle.
chart_period (str, optional) – Chart reporting period.
chart_source (str, optional) – Chart data source.
**fig_kw (Any) – Chart figure kwargs (currently unused).
- Raises:
AggregationColumnError – Missing agg_column for a non-count aggregation, a non-numeric agg_column for a non-count aggregation, or an agg_column named ‘ring’ or ‘wedge’.
AggregationFunctionError – Unexpected aggregation function value.
EmptyDataFrameError – Unexpected empty DataFrame.
KeyError – date_column or agg_column not in DataFrame.
MissingDatetimeError – Unexpected data[date_column] dtype, or data[date_column] contains NaT values.
ModeError – Unexpected mode value is passed.
- Returns:
A tuple containing a DataFrame with the aggregate values used to create the chart, the matplotlib chart Figure and Axes objects.
- Return type:
tuple[DataFrame, Figure, Axes]
Utility
- dataclocklib.utility.add_colorbar(ax, fig, cmap_name, cmap_reverse, vmax, dtype=<class 'numpy.float64'>, vmin=1)[source]
Add a colorbar to a figure, sharing the provided axis.
Values below vmin and NaN values are mapped to white.
- Parameters:
ax (Axes) – Chart Axis.
fig (Figure) – Chart Figure.
cmap_name (str) – Name of matplotlib/PyPalettes colormap.
cmap_reverse (bool) – Reverse cmap colors flag.
vmax (float) – Maximum value of the colorbar.
dtype (DTypeLike, optional) – Data type for colorbar values.
vmin (float, optional) – Minimum value of the colorbar.
- Returns:
A Colorbar object with a cmap and normalised cmap.
- Return type:
Colorbar
- dataclocklib.utility.add_text(ax, x, y, text=None, **kwargs)[source]
Annotate a position on an axis denoted by xy with text.
- Parameters:
ax (Axes) – Axis to annotate.
x (float) – Axis x position.
y (float) – Axis y position.
text (str, optional) – Text to annotate; an empty string if None.
**kwargs (Any) – Text properties passed to Axes.text.
- Returns:
Text object with annotation.
- Return type:
Text
- dataclocklib.utility.add_wedge_labels(ax, font_scale_factor, ring_scale_factor, ring_text_spacing, max_radius, theta, width, wedge_labels)[source]
Add scaled and rotated labels around each data clock wedge.
Labels are placed using Axes.text to facilitate custom rotation of the text, which is based on the angle of the wedge being annotated. The text is scaled based on the size of the chart Figure and padded away from the polar axis based on the number of rings in the chart.
- Parameters:
ax (Axes) – Chart Axis.
font_scale_factor (float) – Scale factor based on current figure size.
ring_scale_factor (float) – Scale factor based on number of rings.
ring_text_spacing (float) – Text label distance from polar axis.
max_radius (int) – Maximum radius (unique rings + 1).
theta (NDArray[np.float64]) – Angles (radians) for each wedge.
width (float) – Width of each wedge (2 * Pi / number of wedges).
wedge_labels (Sequence[str]) – Label text for each wedge.
- Returns:
None
- Return type:
None
- dataclocklib.utility.assign_temporal_columns(data, date_column, mode)[source]
Assign ring & wedge columns to a DataFrame based on mode.
The mode value is mapped to a predetermined division of a larger unit of time into rings, which are then subdivided by a smaller unit of time into wedges, creating a set of temporal bins. These bins are assigned as ‘ring’ and ‘wedge’ columns.
‘YEAR_WEEK’ rings are calendar years, with ISO week numbers clamped so that weeks never cross a calendar year boundary; week 1 therefore spans 4 - 10 days and week 52 spans 5 - 12 days. ‘WEEK_DAY’ rings are ISO year-weeks (YYYYWW), so a calendar-year filter can include a partial ISO week from a neighbouring year (e.g. 2010-01-01 is in ring 200953).
- Parameters:
data (DataFrame) – DataFrame containing data to visualise.
date_column (str) – Name of DataFrame datetime64 column.
mode (Mode, optional) – A mode key representing the temporal bins used in the chart; ‘YEAR_MONTH’, ‘YEAR_WEEK’, ‘WEEK_DAY’, ‘DOW_HOUR’ & ‘DAY_HOUR’.
- Raises:
ModeError – Unexpected mode value is passed.
- Returns:
A DataFrame with ‘ring’ & ‘wedge’ columns assigned.
- Return type:
DataFrame
- dataclocklib.utility.aggregate_temporal_columns(data, agg_column, agg, mode)[source]
Aggregate values in agg_column using pass aggregate function.
Groups the DataFrame by the temporal ‘ring’ and ‘wedge’ columns, before applying the aggregate function to the chosen aggregation column. Missing ring/wedge combinations are filled with 0.
NOTE: The ‘ring’ & ‘wedge’ columns are assigned by the utility function assign_temporal_columns.
- Parameters:
data (DataFrame) – DataFrame containing data to aggregate.
agg_column (str) – DataFrame Column to aggregate.
agg (Aggregation) – Aggregation function; ‘count’, ‘max’, ‘mean’, ‘median’, ‘min’ & ‘sum’.
mode (Mode) – A mode key representing the temporal bins used in the chart; ‘YEAR_MONTH’, ‘YEAR_WEEK’, ‘WEEK_DAY’, ‘DOW_HOUR’ & ‘DAY_HOUR’.
- Raises:
ModeError – Unexpected mode value is passed.
ValueError – Missing ‘ring’ & ‘wedge’ columns.
- Returns:
A DataFrame with aggregate values in a new column named after the aggregate function.
- Return type:
DataFrame
- dataclocklib.utility.get_figure_dimensions(wedges)[source]
Calculate an optimal data clock figure size based on wedge count.
For most data clock charts, a minimum of 0.70 inches of figure space per wedge appears to work best. The best figure shape for this type of chart is square, given the circular nature of the chart.
NOTE: The minimum figure size is capped at (10.0, 10.0).
Example
>>> get_figure_dimensions(168) (11.0, 11.0)
- Parameters:
wedges (int) – Number of wedges (number of rings * wedges per ring).
- Returns:
A tuple containing the height & width of the square figure in inches.
- Return type:
tuple[float, float]
Exceptions
- class dataclocklib.exceptions.AggregationColumnError(agg, reason=None)[source]
Raised on a missing or unsuitable aggregation column.
- Parameters:
agg (str)
reason (str | None)
- class dataclocklib.exceptions.AggregationFunctionError(agg, valid_functions)[source]
Raised on unexpected aggregation function.
- Parameters:
agg (str)
valid_functions (Iterable[str])
- class dataclocklib.exceptions.EmptyDataFrameError(data)[source]
Raised on empty DataFrame.
- Parameters:
data (DataFrame)