Check out mosaic from uwdata which works on top of Observable plot
Or plotly-resampler which works on top of plotly and uses the rust package tsdownsample to aggregate on the 4pixels per pixel shown level (to make antialias work)
the grammar of graphics approach really is a great abstraction, and I'd love to see xy work in that direction
Thanks! XY already uses a similar pixel-aware approach
long line and area traces are reduced in Rust using M4 to produce viewportsized extrema, that is then refined as you zoom. Dense scatter plots use a fixed-size density grid plus a representative sample
I’m not convinced GPU acceleration is a meaningful advantage for most charting use cases. Most dashboards don’t render enough data for it to matter. Once a chart is dense enough for rendering to become the bottleneck, it normally is already be too crowded to be meaningful.
Zooming can justify supporting larger datasets, but sampling/viewport culling and level of detail often avoid drawing unnecessary points...
I think meaningful is in the eye of the beholder. The library is designed such that the trace buffers are directly used as inputs to the WebGL2 drawing contexts to avoid unnecessary copying throughout the stack, which does make a difference when rendering on mobile and embedded devices with limited CPU but often having GPU resources available.
It depends on how much data you are planning on showing, but as you can see from the benchmarks its also more performant than other python charting libs for small data.
We also built this library for extreme customization with CSS/Tailwind support so rendering large amounts of data is an important but not the only advantage.
I’m not trying to justify whether I'd personally use it, I was simply raising the question because the tradeoffs are interesting, and the replies, including Evidlo’s, have already taught me something.
I tried the library before commenting, it’s a cool project. The performance improvements at larger scales I've found are real. I was mainly just trying to have a conversation to learn where we all learn!
“Just move on” seems like a curious response on a discussion forum, though :)
Interesting approach to large scale visualisation. Moving reduction into Rust and sending screen bounded data to WebGL seems much more sensible than pushing millions of raw points into the browser. How does it perform with real time updates? I am assuming it is much more performant? Any plans for a prod deployment?
Yes, it’s significantly faster than existing Python charting libraries for real-time updates.
Instead of serializing and sending the full dataset as JSON, it sends compact typed binary buffers and only the screen-relevant data reducing payload size and browser-side work.
Interesting; how do the examples compare to datashader?
Edit: for my use cases, I use napari (~1e7-8 points) if I need true interactivity; otherwise, datashader/holoviz, or even just fast-histogram's 2D histograms work.
For extremely large point clouds, these caveats[0] still apply. It irks me when people make dense scatterplots without any indication of just how dense some portions are.
Still, if it can indeed handle 1e10 points, that's pretty impressive.
it's possible to render data out-of-core with XY, allowing it to render the entirety of OpenStreetMaps (that's 10,742,674,832 nodes!) with sub-second pan/zooms. it's a bit difficult to host online but you can try it out locally: https://github.com/reflex-dev/xy/tree/main/examples/osm
Or plotly-resampler which works on top of plotly and uses the rust package tsdownsample to aggregate on the 4pixels per pixel shown level (to make antialias work)
the grammar of graphics approach really is a great abstraction, and I'd love to see xy work in that direction
long line and area traces are reduced in Rust using M4 to produce viewportsized extrema, that is then refined as you zoom. Dense scatter plots use a fixed-size density grid plus a representative sample
We also built this library for extreme customization with CSS/Tailwind support so rendering large amounts of data is an important but not the only advantage.
You don't have to justify your decision to people here, literally just move on with your life and forget about it.
I tried the library before commenting, it’s a cool project. The performance improvements at larger scales I've found are real. I was mainly just trying to have a conversation to learn where we all learn!
“Just move on” seems like a curious response on a discussion forum, though :)
Instead of serializing and sending the full dataset as JSON, it sends compact typed binary buffers and only the screen-relevant data reducing payload size and browser-side work.
More detail here https://github.com/reflex-dev/xy#how-it-works
Edit: for my use cases, I use napari (~1e7-8 points) if I need true interactivity; otherwise, datashader/holoviz, or even just fast-histogram's 2D histograms work.
For extremely large point clouds, these caveats[0] still apply. It irks me when people make dense scatterplots without any indication of just how dense some portions are.
Still, if it can indeed handle 1e10 points, that's pretty impressive.
[0]: https://datashader.org/user_guide/Plotting_Pitfalls.html
Love rust as the impl.
> XY is an extremely fast, interactive, customizable Python charting library
which is it?