# \#query

**URL:** https://discourse.julialang.org/tag/query/120.md

[Latest](https://discourse.julialang.org/latest.md) · [Categories](https://discourse.julialang.org/categories.md) · [Tags](https://discourse.julialang.org/tags.md)

---

## [\[ANN\] SQLCollections.jl – use Julia data manipulation functions for databases](https://discourse.julialang.org/t/ann-sqlcollections-jl-use-julia-data-manipulation-functions-for-databases/119244)

<div class="topic-metadata">

**Author:** [@aplavin](https://discourse.julialang.org/u/aplavin)\
**Replies:** 31\
**Last updated:** [October 28, 2024, 6:01pm UTC](https://discourse.julialang.org/t/ann-sqlcollections-jl-use-julia-data-manipulation-functions-for-databases/119244 "2024-10-28T18:01:39Z")

</div>

Imagine if you could use the exact same code to manipulate regular Julia collections and SQL tables… SQLCollections.jl is the package to achieve that! There is no shortage of Julia packages that provide convenient int…

---

## [Basic function usage in Query @filter](https://discourse.julialang.org/t/basic-function-usage-in-query-filter/114642)

<div class="topic-metadata">

**Author:** [@seadoo](https://discourse.julialang.org/u/seadoo)\
**Replies:** 9\
**Last updated:** [May 23, 2024, 9:04pm UTC](https://discourse.julialang.org/t/basic-function-usage-in-query-filter/114642 "2024-05-23T21:04:00Z")

</div>

Is the @filter macro from Query.jl only useful for simple things like \_.a \> 1 and so on? I’ve tried to use ismissing which just returns a boolean and it does not return anything (but also does not return any errors). Br…

---

## [Help improving the speed of a DataFrames operation](https://discourse.julialang.org/t/help-improving-the-speed-of-a-dataframes-operation/107615)

<div class="topic-metadata">

**Author:** [@abelsiqueira](https://discourse.julialang.org/u/abelsiqueira)\
**Replies:** 37\
**Last updated:** [February 22, 2024, 9:25am UTC](https://discourse.julialang.org/t/help-improving-the-speed-of-a-dataframes-operation/107615 "2024-02-22T09:25:18Z")

</div>

Hi all, quick question (hopefully) related to doing an operation on DataFrames. I am trying to improve the transformation below so any help would be great. More context and MWE are at the bottom. Thanks for all the hel…

---

## [Dynamic selection of columns Query.jl](https://discourse.julialang.org/t/dynamic-selection-of-columns-query-jl/79526)

<div class="topic-metadata">

**Author:** [@davide-f](https://discourse.julialang.org/u/davide-f)\
**Replies:** 4\
**Last updated:** [October 16, 2023, 2:56pm UTC](https://discourse.julialang.org/t/dynamic-selection-of-columns-query-jl/79526 "2023-10-16T14:56:25Z")

</div>

Dear Community, I am trying to dynamically select columns using Query.jl like in the following script, but I am unable to do so. Any ideas? Thank you list\_cols = \[“col1”, “col2”…\] select = df |\> @select(∈(list\_cols)…

---

## [Filtering DataFrame on variable attribute type](https://discourse.julialang.org/t/filtering-dataframe-on-variable-attribute-type/104059)

<div class="topic-metadata">

**Author:** [@edo-0071](https://discourse.julialang.org/u/edo-0071)\
**Replies:** 1\
**Last updated:** [September 20, 2023, 5:49pm UTC](https://discourse.julialang.org/t/filtering-dataframe-on-variable-attribute-type/104059 "2023-09-20T17:49:50Z")

</div>

Hi, i have a simple DataFrame and want to extract rows that match certain values ​​for some (not necessarily all) attributes. Starting from here, res = @from i in df begin @where i.A = "a1" @select i @collect DataFrame…

---

## [DataFrame sort Performance using Query.jl vs SAS PROC SORT](https://discourse.julialang.org/t/dataframe-sort-performance-using-query-jl-vs-sas-proc-sort/41225)

<div class="topic-metadata">

**Author:** [@Dominic\_Pazzula](https://discourse.julialang.org/u/Dominic_Pazzula)\
**Replies:** 19\
**Last updated:** [August 11, 2023, 9:14pm UTC](https://discourse.julialang.org/t/dataframe-sort-performance-using-query-jl-vs-sas-proc-sort/41225 "2023-08-11T21:14:38Z")

</div>

I’m still learning Julia and am converting some SAS code. I have a large DataFrame (3.3M rows, 6 columns) that I am trying to sort. The performance compared to SAS is abysmal. function sort1(simstates) simstates2…

---

## [Query - column names with spaces](https://discourse.julialang.org/t/query-column-names-with-spaces/56232)

<div class="topic-metadata">

**Author:** [@statspy](https://discourse.julialang.org/u/statspy)\
**Replies:** 5\
**Last updated:** [April 6, 2023, 8:11am UTC](https://discourse.julialang.org/t/query-column-names-with-spaces/56232 "2023-04-06T08:11:15Z")

</div>

I have this df: d = DataFrame(Symbol("Full Name") =\> \["Mike","Paul","Andres"\], Symbol("Age") =\> 10:12) so far so good. Now i want to do this Query: using Query @from i in d begin @where i.Age == 11 @select {i."Full Name…

---

## [Lifting function in Query.jl](https://discourse.julialang.org/t/lifting-function-in-query-jl/96599)

<div class="topic-metadata">

**Author:** [@myrddin89](https://discourse.julialang.org/u/myrddin89)\
**Replies:** 2\
**Last updated:** [March 25, 2023, 6:17pm UTC](https://discourse.julialang.org/t/lifting-function-in-query-jl/96599 "2023-03-25T18:17:19Z")

</div>

I’m having troubles using the lifting functionality in Query.jl and DataValues.jl with dot syntax. If I simply try the following using DataValues a = DataValue(0.4) a .\* 2 I get the this error: ERROR: MethodError: no…

---

## [Query.jl fails handle CSV.Rows](https://discourse.julialang.org/t/query-jl-fails-handle-csv-rows/95063)

<div class="topic-metadata">

**Author:** [@purplesabbath](https://discourse.julialang.org/u/purplesabbath)\
**Replies:** 1\
**Last updated:** [February 23, 2023, 8:17am UTC](https://discourse.julialang.org/t/query-jl-fails-handle-csv-rows/95063 "2023-02-23T08:17:06Z")

</div>

I got an dataset too large for memory, I want to combine CSV.Rows and Query.jl to handle data, but after @groupby, I can’t use @map to summarize data. using DataFrames using CSV using Query data\_path = "D:\\\\chen\\\\det…

---

## [Summarizing Float64 values after grouping by different time scales from DateTime object](https://discourse.julialang.org/t/summarizing-float64-values-after-grouping-by-different-time-scales-from-datetime-object/90320)

<div class="topic-metadata">

**Author:** [@Hugo](https://discourse.julialang.org/u/Hugo)\
**Replies:** 2\
**Last updated:** [December 21, 2022, 3:56pm UTC](https://discourse.julialang.org/t/summarizing-float64-values-after-grouping-by-different-time-scales-from-datetime-object/90320 "2022-12-21T15:56:59Z")

</div>

I have a data frame with two columns. Column time has datetime elements while column values has Float64 elements. I want to group the column time by year, month, day, and hour to calculate the average of the values. I c…

---

## [How to query with multiple conditions with Query.jl?](https://discourse.julialang.org/t/how-to-query-with-multiple-conditions-with-query-jl/78240)

<div class="topic-metadata">

**Author:** [@SubTer](https://discourse.julialang.org/u/SubTer)\
**Replies:** 2\
**Last updated:** [March 22, 2022, 12:50pm UTC](https://discourse.julialang.org/t/how-to-query-with-multiple-conditions-with-query-jl/78240 "2022-03-22T12:50:35Z")

</div>

hey, hoping someone can help me out with the syntax Julia Query uses, or maybe I am missing something, but basically trying to replicate a query command using contains, instead of having bunch of “||” statements. examp…

---

## [ACOPF Problem](https://discourse.julialang.org/t/acopf-problem/72375)

<div class="topic-metadata">

**Author:** [@priod](https://discourse.julialang.org/u/priod)\
**Replies:** 2\
**Last updated:** [December 1, 2021, 7:40pm UTC](https://discourse.julialang.org/t/acopf-problem/72375 "2021-12-01T19:40:38Z")

</div>

I am trying to run an ACOPF (Alternating current optimal power flow) in the two-node network for a variable load in time “t”. The result of the model is just the zero values and nothing else. Can someone suggest to me a …

---

## [Use the non-missing elements of DataFrame A to update DataFrame B](https://discourse.julialang.org/t/use-the-non-missing-elements-of-dataframe-a-to-update-dataframe-b/64658)

<div class="topic-metadata">

**Author:** [@fipelle](https://discourse.julialang.org/u/fipelle)\
**Replies:** 9\
**Last updated:** [July 18, 2021, 8:53am UTC](https://discourse.julialang.org/t/use-the-non-missing-elements-of-dataframe-a-to-update-dataframe-b/64658 "2021-07-18T08:53:57Z")

</div>

Hi, I am trying to find an elegant solution to update a DataFrame (B), using the non-missing elements in another DataFrame (A) with the same columns. For instance, with using Dates, DataFrames; # Dates date\_1 = Date("…

---

## [OHDSICohortExpressions.jl](https://discourse.julialang.org/t/ohdsicohortexpressions-jl/62525)

<div class="topic-metadata">

**Author:** [@cce](https://discourse.julialang.org/u/cce)\
**Replies:** 1\
**Last updated:** [June 7, 2021, 3:07pm UTC](https://discourse.julialang.org/t/ohdsicohortexpressions-jl/62525 "2021-06-07T15:07:57Z")

</div>

Observational Health Data Sciences and Informatics (or OHDSI) is a program to bring out the value of health data through large-scale analytics. OHDSI defines the OMOP CDM, a data warehouse database schema for storing …

---

## [Data manipulation using dataframes](https://discourse.julialang.org/t/data-manipulation-using-dataframes/57972)

<div class="topic-metadata">

**Author:** [@mjanun](https://discourse.julialang.org/u/mjanun)\
**Replies:** 7\
**Last updated:** [March 27, 2021, 1:07pm UTC](https://discourse.julialang.org/t/data-manipulation-using-dataframes/57972 "2021-03-27T13:07:39Z")

</div>

Assume that a dataframe is given as: df\_a = DataFrame(Categ=\["Alpa","Beta","Gamma","Omega","Alpa","Beta","Gamma","Omega"\], T1 =rand(8),T2=rand(8),T3 =rand(8),T4 =rand(8),T5 =rand(8),T6 =rand(8),T7 =rand(8),T8 =rand(8),T…

---

## [\[Query.jl\] Way to map many variables to a single function](https://discourse.julialang.org/t/query-jl-way-to-map-many-variables-to-a-single-function/56675)

<div class="topic-metadata">

**Author:** [@Storopoli](https://discourse.julialang.org/u/Storopoli)\
**Replies:** 0\
**Last updated:** [March 7, 2021, 1:29pm UTC](https://discourse.julialang.org/t/query-jl-way-to-map-many-variables-to-a-single-function/56675 "2021-03-07T13:29:50Z")

</div>

Here is the code in DataFrames: @chain DataFrame(tbl) begin dropmissing(\[:NT\_GER, :NT\_FG, :NT\_CE\]) groupby(\[:TP\_SEXO, :CO\_TURNO\_GRADUACAO\], skipmissing=true) combine(nrow, \[:NT\_GER, :NT\_FG, :NT\_C…

---

## [Collect a singleton query result as a dict with no missing values](https://discourse.julialang.org/t/collect-a-singleton-query-result-as-a-dict-with-no-missing-values/49157)

<div class="topic-metadata">

**Author:** [@mkarikom](https://discourse.julialang.org/u/mkarikom)\
**Replies:** 0\
**Last updated:** [October 28, 2020, 1:14am UTC](https://discourse.julialang.org/t/collect-a-singleton-query-result-as-a-dict-with-no-missing-values/49157 "2020-10-28T01:14:09Z")

</div>

How can I refactor the following queries so that the tests pass? df = DataFrame(:a=\>\[1,2\],:b=\>\["a","b"\],:c=\>\[1,missing\]) 2×3 DataFrame │ Row │ a │ b │ c │ │ │ Int64 │ String │ Int64? │ ├─────┼───────…

---

## [Join Dataframes on different IDs within date range](https://discourse.julialang.org/t/join-dataframes-on-different-ids-within-date-range/36425)

<div class="topic-metadata">

**Author:** [@mp-crypto](https://discourse.julialang.org/u/mp-crypto)\
**Replies:** 1\
**Last updated:** [October 4, 2020, 2:05am UTC](https://discourse.julialang.org/t/join-dataframes-on-different-ids-within-date-range/36425 "2020-10-04T02:05:42Z")

</div>

Hello, I am trying to merge two datasets. I am using query (as in the code below), but I should merge on i.ticker NOT equals j.ticker. Basically, I should merge to a focal i.ticker all other j.tickers whose i.datestart\<j…

---

## [LINQ inner joins and select all in Query.jl](https://discourse.julialang.org/t/linq-inner-joins-and-select-all-in-query-jl/46556)

<div class="topic-metadata">

**Author:** [@mkarikom](https://discourse.julialang.org/u/mkarikom)\
**Replies:** 7\
**Last updated:** [September 14, 2020, 2:59pm UTC](https://discourse.julialang.org/t/linq-inner-joins-and-select-all-in-query-jl/46556 "2020-09-14T14:59:43Z")

</div>

This MWE is supposed to do an inner join (i.a equals j.c || i.b equals j.e) and output all the elements of the first df: df1 = DataFrame(a=\[1,2,3\], b=\[1.,2.,3.\]) df2 = DataFrame(c=\[2,4,2\], d=\["John", "Jim","Sally"\],e=…

---

## [Does it make sense for @orderby and @groupby not to commute?](https://discourse.julialang.org/t/does-it-make-sense-for-orderby-and-groupby-not-to-commute/45920)

<div class="topic-metadata">

**Author:** [@conditionality](https://discourse.julialang.org/u/conditionality)\
**Replies:** 2\
**Last updated:** [September 2, 2020, 6:05pm UTC](https://discourse.julialang.org/t/does-it-make-sense-for-orderby-and-groupby-not-to-commute/45920 "2020-09-02T18:05:39Z")

</div>

I have a dataframe which I want to group by some column value, and within each group I want the rows sorted by another column. I can do an @orderby |\> @groupby, however I feel like a @groupby |\> @orderby should make equa…

---

## [@replacena not working with variables as replacement value](https://discourse.julialang.org/t/replacena-not-working-with-variables-as-replacement-value/45401)

<div class="topic-metadata">

**Author:** [@TheRoniOne](https://discourse.julialang.org/u/TheRoniOne)\
**Replies:** 4\
**Last updated:** [August 23, 2020, 5:18pm UTC](https://discourse.julialang.org/t/replacena-not-working-with-variables-as-replacement-value/45401 "2020-08-23T17:18:25Z")

</div>

Hi, was trying to practice some simple data cleaning with Query.jl modifying a bit an example of the documentation. What I intended to do was to replace all the missing values in a column with the mean of that column. …

---

## [Stack and unstack in Query.jl](https://discourse.julialang.org/t/stack-and-unstack-in-query-jl/45381)

<div class="topic-metadata">

**Author:** [@tnederlof](https://discourse.julialang.org/u/tnederlof)\
**Replies:** 5\
**Last updated:** [August 23, 2020, 2:55pm UTC](https://discourse.julialang.org/t/stack-and-unstack-in-query-jl/45381 "2020-08-23T14:55:14Z")

</div>

Something I used a lot when I am in the R/Tidyverse world is spread and gather from tidyr most of which I can replicate using stack and unstack in DataFrames.jl. Within Query.jl is there a way to accomplish this? If not …

---

## [Query @select nested in a function](https://discourse.julialang.org/t/query-select-nested-in-a-function/45218)

<div class="topic-metadata">

**Author:** [@elalaouifaris](https://discourse.julialang.org/u/elalaouifaris)\
**Replies:** 0\
**Last updated:** [August 19, 2020, 3:53pm UTC](https://discourse.julialang.org/t/query-select-nested-in-a-function/45218 "2020-08-19T15:53:38Z")

</div>

Hi, I have an issue using Query @select using a function call parameter to select columns: using Query, VegaDatasets iris = dataset("iris") iris |\> @select(startswith("petal")) # This works of course function get\_da…

---

## [Count cumulative number of unique elements](https://discourse.julialang.org/t/count-cumulative-number-of-unique-elements/42736)

<div class="topic-metadata">

**Author:** [@tpoisot](https://discourse.julialang.org/u/tpoisot)\
**Replies:** 1\
**Last updated:** [July 8, 2020, 3:10pm UTC](https://discourse.julialang.org/t/count-cumulative-number-of-unique-elements/42736 "2020-07-08T15:10:45Z")

</div>

I am trying to produce a species discovery curve, which is basically the number of unique species observed over time. So, at time t, the value is length(unique(\_.species\_name)) for all observations up to time t - this in…

---

## [Query - missing vs. isna](https://discourse.julialang.org/t/query-missing-vs-isna/39952)

<div class="topic-metadata">

**Author:** [@stej](https://discourse.julialang.org/u/stej)\
**Replies:** 7\
**Last updated:** [May 22, 2020, 2:00pm UTC](https://discourse.julialang.org/t/query-missing-vs-isna/39952 "2020-05-22T14:00:24Z")

</div>

Hi all, I have a question about ismissing and isna. I found Query.jl - filtering on missing data so I started using that. I thought that isna is really needed whenever I use Query.jl. But that’s not the case. Example …

---

## [Query @mutate changes other column type](https://discourse.julialang.org/t/query-mutate-changes-other-column-type/39292)

<div class="topic-metadata">

**Author:** [@stej](https://discourse.julialang.org/u/stej)\
**Replies:** 8\
**Last updated:** [May 13, 2020, 5:39pm UTC](https://discourse.julialang.org/t/query-mutate-changes-other-column-type/39292 "2020-05-13T17:39:50Z")

</div>

Hi all, this is my smallest code that I fight with: contents = """ "5674012","aa66aa66" "5674012","9b4e08e5" "5674012","b036aa66,b036aa67,b036aa68" """; batches = CSV.File(IOBuffer(contents); header = \["X1", "Splits"\],…

---

## [Converting DataValues.DataValue{String} to String](https://discourse.julialang.org/t/converting-datavalues-datavalue-string-to-string/38941)

<div class="topic-metadata">

**Author:** [@stej](https://discourse.julialang.org/u/stej)\
**Replies:** 10\
**Last updated:** [May 11, 2020, 3:51pm UTC](https://discourse.julialang.org/t/converting-datavalues-datavalue-string-to-string/38941 "2020-05-11T15:51:45Z")

</div>

Hi all, just continuing from Read CSV and change rows later - #6 by pdeffebach I’m trying to convert a column like this: using CSV using DataFrames using Query file = """ "X1", "X2", "X3", "x4", "Splits" "5674012","53…

---

## [Issues querying a DataFrame](https://discourse.julialang.org/t/issues-querying-a-dataframe/34943)

<div class="topic-metadata">

**Author:** [@bensetterholm](https://discourse.julialang.org/u/bensetterholm)\
**Replies:** 5\
**Last updated:** [February 21, 2020, 8:21pm UTC](https://discourse.julialang.org/t/issues-querying-a-dataframe/34943 "2020-02-21T20:21:57Z")

</div>

Suppose I have a DataFrame df which was generated by a Dictionary which may include whitespace in the keys using DataFrames myDict = Dict() myDict\["aKey"\] = 1:10 myDict\["anotherKey"\] = zeros(10) myDict\["a\_nice\_key"\] = …

---

## [Can't use NDSparse Created Using Query.jl](https://discourse.julialang.org/t/cant-use-ndsparse-created-using-query-jl/33620)

<div class="topic-metadata">

**Author:** [@HBreddam](https://discourse.julialang.org/u/HBreddam)\
**Replies:** 2\
**Last updated:** [January 22, 2020, 3:25pm UTC](https://discourse.julialang.org/t/cant-use-ndsparse-created-using-query-jl/33620 "2020-01-22T15:25:22Z")

</div>

Hi I want i have my data in Indexed tables and want to use Query.jl to make NDSparse tables for use in a creating a MIP model using JuMP. However when i have created the NDSparse test as seen below, i can’t get the dat…

---

## [Queryverse: How to @select columns with specific eltype? (using \`oftype\` function)](https://discourse.julialang.org/t/queryverse-how-to-select-columns-with-specific-eltype-using-oftype-function/28709)

<div class="topic-metadata">

**Author:** [@Anselm\_Jeong](https://discourse.julialang.org/u/Anselm_Jeong)\
**Replies:** 3\
**Last updated:** [September 14, 2019, 10:14am UTC](https://discourse.julialang.org/t/queryverse-how-to-select-columns-with-specific-eltype-using-oftype-function/28709 "2019-09-14T10:14:47Z")

</div>

Dear all Queryverse users. In an auxilary package QueryOperators.NamedTupleUtilities, there is a helper function named oftype. I guess it is related with selecting columns with specific eltype. But I cannot figure out h…

[Next page](https://discourse.julialang.org/tag/query/120.md?match_all_tags=true&page=1&tags%5B%5D=query)
