feat: add DataFrameLike parameter for cross-backend dataframe inputs - #1144
feat: add DataFrameLike parameter for cross-backend dataframe inputs#1144ghostiee-11 wants to merge 6 commits into
Conversation
`param.DataFrame` is restricted to pandas. `DataFrameLike` accepts any object Narwhals recognises (pandas, Polars, PyArrow, cuDF, Modin) and passes it through unchanged, so existing pandas-only code is unaffected (`param.DataFrame` is not touched). * New `DataFrameLike(ClassSelector)` validating via `narwhals.from_native(eager_only=not allow_lazy, pass_through=False)`. Narwhals is an optional dependency, deferred like pandas is for `DataFrame`. * Same `rows` / `columns` / `ordered` slots as `DataFrame`, driven through the Narwhals wrapper so they work on every backend. Column names read via `collect_schema().names()` so lazy frames are not implicitly collected. * `allow_lazy=True` opts into lazy frames (Polars LazyFrame, Dask, DuckDB); row-count validation is skipped for lazy frames. * Backend-neutral `serialize` (list of records via Narwhals); `deserialize` reuses `DataFrame.deserialize` since JSON carries no backend information. * `_length_bounds_check` extracted to a module-level helper shared by `DataFrame` and `DataFrameLike` (behaviour-preserving; testpandas unchanged). * tests/testdataframelike.py covering pandas / Polars / PyArrow / lazy / serialization; narwhals + polars added to test-only dependencies.
* Raise a clear ImportError naming the install command when the optional narwhals package is missing, instead of a bare ModuleNotFoundError (declaration-time fail-fast, matching how DataFrame fails on missing pandas). * Document the serialization asymmetry (backend-neutral records out, pandas in) and that cuDF/Modin are Narwhals-supported but not run in CI (cuDF is GPU-only, Modin's pinned deps conflict with the test environment). * Annotate the inherited in-place ordered defaulting as deliberate DataFrame parity. * Add a skip-guarded Modin test and add narwhals + polars to the type-check environment so pyright validates the Narwhals API rather than skipping an unresolved import.
Remove cuDF/Modin name-drops from the docstring, error message and tests. They are reachable through Narwhals like any other backend but are not exercised here (no GPU; Modin's pinned deps conflict), so naming them as features overclaims. The validation path is described generically as "any Narwhals-supported backend" with pandas, Polars and PyArrow as the tested set. Drops the permanently-skipped Modin test.
|
Hey @philippjfr !! Looking for your views over this. Thanks!! |
hoxbro
left a comment
There was a problem hiding this comment.
Left some comments.
I'm not entirely sure if we should keep the underlying DataFrame or convert it to a narwhals DataFrame/LazyFrame.
A user would need to do the conversion in every Parameterized class methods which uses it as they API is very different for the DataFrame APIs.
Also how much AI have you used for this PR?
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #1144 +/- ##
==========================================
+ Coverage 86.75% 86.80% +0.04%
==========================================
Files 9 9
Lines 5302 5380 +78
==========================================
+ Hits 4600 4670 +70
- Misses 702 710 +8 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Thanks for the review. On native vs. narwhals return: pass-through is intentional. Returning a narwhals.DataFrame would break any consumer that calls On AI usage: drafting and refactoring assistant. Design is mine and matches what I argued for in #975. Will trim the LLM-flavoured prose in the docstring this round. |
* Move _get_narwhals to _utils and use narwhals.stable.v2 * Rename allow_lazy to eager_only (default True) * Validate row count on LazyFrame via narwhals .count() instead of skipping * Skip collect_schema() unless columns/ordered or a lazy row check needs it * Compact docstring, fix narwhals URL, drop noisy inline comment * Rewrite tests as plain pytest functions, drop PARAM_TEST_NARWHALS env var, scope lazy tests to polars
…eckers * tests/testdefaults.py: append DataFrameLike to skip list when narwhals is unavailable, matching the existing pandas/numpy pattern * param/parameters.py: restructure DataFrameLike._validate so cols and schema have non-Optional types in the branches that use them, fixing pyrefly/pyright/ty errors flagged on CI
|
Hey @philippjfr, can you review the changes now. Thankyou!! |
|
Apologies I missed your message. I will attempt to get this merged once we get the latest patch release out. |
|
No worries :) Thankyou @philippjfr |
param.DataFrameonly acceptspandas.DataFrame, so there is no way to declare a parameter that holds tabular data when the value might be Polars, PyArrow, or another backend. This adds a newDataFrameLikeparameter that validates anything the Narwhals protocol recognises and passes the native object through unchanged.param.DataFrameis deliberately left untouched, so existing pandas-only code keeps its guarantee. This is the separate-class direction discussed in #975; serialization backend-preservation is intentionally left as an open question there.Same
rows/columns/orderedslots asDataFrame(driven through Narwhals so they work on every backend), plusallow_lazy=Truefor PolarsLazyFrame/ Dask / DuckDB with no implicit collect. Narwhals is an optional dependency, deferred like pandas is forDataFrame, with a clear install message if missing.Before


After
Tested with pandas, Polars (eager + lazy), and PyArrow; full suite 1550 passed,
testpandas.pyunchanged. Validation goes entirely through Narwhals, so any other Narwhals-supported backend uses the identical code path.