# Can Sagemaker Julia query a S3/Athena table with SQL?

**URL:** <https://discourse.julialang.org/t/can-sagemaker-julia-query-a-s3-athena-table-with-sql/66257>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [August 12, 2021, 10:38am UTC](https://discourse.julialang.org/t/can-sagemaker-julia-query-a-s3-athena-table-with-sql/66257 "2021-08-12T10:38:31Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Billpete002](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/billpete002/32/35091_2.png) [@Billpete002](https://discourse.julialang.org/u/Billpete002)\
**Post date:** [August 12, 2021, 10:38am UTC](https://discourse.julialang.org/t/can-sagemaker-julia-query-a-s3-athena-table-with-sql/66257/1 "2021-08-12T10:38:31Z")

</div>

Is it possible run a SQL-esque script to pull data within Julia from AWS’s S3/Athena? My Julia instance is located on AWS Sagemaker Jupyter Notebook.

Using python in SageMaker easily enough:

```julia
import boto3
region = boto3.Session().region_name

from pyathena import connect
import pandas as pd

conn = connect(s3_staging_dir='s3://sagemaker-examplebucket/',
               region_name = region)

df = pd.read_sql("""SELECT 
 something1,
 something2
FROM "customer_data"."sagemaker_data"
WHERE 
    something = 0;""", conn)

```

For my instance of SageMaker Notebook using Julia I can connect and put/get an existing file, but I can’t find anything that allows new queries like Python’s pyathena library:

```julia
using AWS, AWSS3, Serialization
struct SampleData
  a::Int
  b::String
end

d=SampleData(1,"sss")
aws = global_aws_config(; region="us-west-2")
b = IOBuffer()
serialize(b, d)

s3_put(aws, "sagemaker-examplebucket","myfile.bin", b.data)

ddat = s3_get(aws, "sagemaker-examplebucket","myfile.bin")
d2 = deserialize(IOBuffer(ddat))

@assert d == d2

```

---

<div class="post-metadata">

**Author:** ![propelledanalytics](https://avatars.discourse-cdn.com/v4/letter/p/e47774/32.png) [@propelledanalytics](https://discourse.julialang.org/u/propelledanalytics)\
**Post date:** [August 12, 2021, 9:14pm UTC](https://discourse.julialang.org/t/can-sagemaker-julia-query-a-s3-athena-table-with-sql/66257/2 "2021-08-12T21:14:06Z")

</div>

> [@Billpete002](#):
>
> Is it possible run a SQL-esque script to pull data within Julia from AWS’s S3/Athena? My Julia instance is located on AWS Sagemaker Jupyter Notebook.
> 
> Using python in SageMaker easily enough:
> 
> ```julia
> import boto3
> region = boto3.Session().region_name
> 
> from pyathena import connect
> import pandas as pd
> 
> conn = connect(s3_staging_dir='s3://sagemaker-examplebucket/',
> region_name = region)
> 
> df = pd.read_sql("""SELECT 
> something1,
> something2
> FROM "customer_data"."sagemaker_data"
> WHERE 
> something = 0;""", conn)
> 
> ```

Hi Billpete002,

The SparkSQL.jl package enables Julia programs to work with Spark data using SQL. SparkSQL.jl returns results from Apache Spark queries as Julia DataFrames. You can move Julia data to your Spark query too. A common use case for SparkSQL.jl is machine learning. SparkSQL.jl makes it easy to get data from Spark using SQL, do machine learning in Julia, and return data back to Apache Spark. Example syntax:

```julia
JuliaDataFrame = DataFrame(tickers = ["CRM", "IBM"])
onSpark = toSparkDS(sprk, JuliaDataFrame)
createOrReplaceTempView(onSpark, "julia_data")
query = sql(sprk, "SELECT * FROM spark_data WHERE TICKER IN (SELECT * FROM julia_data)")
results = toJuliaDF(query)
describe(results)

```

To learn more, visit the tutorial page and project pages:

Tutorial page:

> **[SparkSQL.jl Blog](https://propelledanalytics.github.io/Tutorials/)**
>
> Tutorials on how to use the SparkSQL.jl Julia Package.

Project page:

> **[GitHub - propelledanalytics/SparkSQL.jl: SparkSQL.jl enables Julia programs...](https://github.com/propelledanalytics/SparkSQL.jl)**
>
> SparkSQL.jl enables Julia programs to work with Apache Spark data using just SQL. - GitHub - propelledanalytics/SparkSQL.jl: SparkSQL.jl enables Julia programs to work with Apache Spark data using ...

---

<div class="post-metadata">

**Author:** ![Billpete002](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/billpete002/32/35091_2.png) [@Billpete002](https://discourse.julialang.org/u/Billpete002)\
**Post date:** [August 13, 2021, 5:20am UTC](https://discourse.julialang.org/t/can-sagemaker-julia-query-a-s3-athena-table-with-sql/66257/3 "2021-08-13T05:20:11Z")

</div>

Thanks for the answer. Finding libraries for Julia has been a challenge! I don’t have a spark instance set up currently but this looks do-able. It would be great to keep this all within Julia like R and Python can.

---

<div class="post-metadata">

**Author:** ![drizk1](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/drizk1/32/208422_2.png) [@drizk1](https://discourse.julialang.org/u/drizk1)\
**Post date:** [May 10, 2024, 3:00am UTC](https://discourse.julialang.org/t/can-sagemaker-julia-query-a-s3-athena-table-with-sql/66257/4 "2024-05-10T03:00:39Z")

</div>

I know this is a few years late, but with TidierDB.jl and AWS.jl, you can run sql queries and collect the data as a dataframe from AWS’s Athena.

[Here is the documentation for connecting and running queries.](https://tidierorg.github.io/TidierDB.jl/latest/examples/generated/UserGuide/athena/) Hope this helps!
