# Read large stream from STDIN

**URL:** <https://discourse.julialang.org/t/read-large-stream-from-stdin/55627>\
**Category:** General Usage\
**Created:** [February 19, 2021, 4:40pm UTC](https://discourse.julialang.org/t/read-large-stream-from-stdin/55627 "2021-02-19T16:40:24Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![tferic](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tferic/32/12916_2.png) [@tferic](https://discourse.julialang.org/u/tferic)\
**Post date:** [February 19, 2021, 4:40pm UTC](https://discourse.julialang.org/t/read-large-stream-from-stdin/55627/1 "2021-02-19T16:40:24Z")

</div>

Hi  
I would like to use Julia to process large logfiles piped into STDIN, in order to create reports.  
The amount of data is too big to fit into RAM, so I need to process every line by line while reading the stream of data (rather than slurping all data into a variable before starting to process data).

Like this:

```julia
cat my_huge_logfile.log | reporting.jl

```

In Perl, I would use something like this:

```julia
while(<STDIN>) {
    # regex matching on current line
    # do some preprocessing
    # remember selected data in hash/array
}
# reporting based on hash/array

```

What is the equivalent or better way do it in Julia?

---

<div class="post-metadata">

**Author:** ![jmert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmert/32/3161_2.png) [@jmert](https://discourse.julialang.org/u/jmert)\
**Post date:** [February 19, 2021, 4:52pm UTC](https://discourse.julialang.org/t/read-large-stream-from-stdin/55627/2 "2021-02-19T16:52:52Z")

</div>

Something like this?

```julia
function main()
    bytes = 0
    while !eof(stdin)
        line = readline(stdin)
        bytes += length(line)
    end
    println(bytes)
end

main()

```

```bash
$ journalctl -b | julia stream.jl 
21571339

```

---

<div class="post-metadata">

**Author:** ![tferic](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tferic/32/12916_2.png) [@tferic](https://discourse.julialang.org/u/tferic)\
**Post date:** [February 19, 2021, 5:16pm UTC](https://discourse.julialang.org/t/read-large-stream-from-stdin/55627/3 "2021-02-19T17:16:08Z")

</div>

Many thanks @jmert  
I have figured out another way. I will try both and see which is faster.

```julia
for line = readlines()
    # Work with variable line here
end

```

---

<div class="post-metadata">

**Author:** ![jmert](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/jmert/32/3161_2.png) [@jmert](https://discourse.julialang.org/u/jmert)\
**Post date:** [February 19, 2021, 5:23pm UTC](https://discourse.julialang.org/t/read-large-stream-from-stdin/55627/4 "2021-02-19T17:23:13Z")

</div>

Looking at the code with `@edit readlines(stdin)`, it looks like it’ll load everything into RAM as a vector of lines. But that inspection also leads to `eachline(stdin)` which I think is more what you were looking for.

[`eachline()`](https://docs.julialang.org/en/v1/base/io-network/#Base.eachline) docs

---

<div class="post-metadata">

**Author:** ![tferic](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tferic/32/12916_2.png) [@tferic](https://discourse.julialang.org/u/tferic)\
**Post date:** [February 19, 2021, 11:55pm UTC](https://discourse.julialang.org/t/read-large-stream-from-stdin/55627/5 "2021-02-19T23:55:11Z")

</div>

> [@jmert](#):
>
> Looking at the code with `@edit readlines(stdin)` , it looks like it’ll load everything into RAM as a vector of lines. But that inspection also leads to `eachline(stdin)` which I think is more what you were looking for.

@jmert Thanks for taking the time. Yes, I can confirm. I tried `readlines()` on a huge logfile, and it is trying to read everything into memory. That’s not what I want.

The following structure seems to work with large data the way I want:

```julia
for line = eachline()
    # Work with variable line here
end

```

Many thanks for your help!  
Toni

---

<div class="post-metadata">

**Author:** ![tferic](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tferic/32/12916_2.png) [@tferic](https://discourse.julialang.org/u/tferic)\
**Post date:** [February 20, 2021, 9:38pm UTC](https://discourse.julialang.org/t/read-large-stream-from-stdin/55627/6 "2021-02-20T21:38:25Z")

</div>

I noticed that there are performance issues with `eachline()` as well.  
For these performance reasons, I have opened a new thread here:

> [@Bad performance of eachline() on STDIN](https://discourse.julialang.org/t/bad-performance-of-eachline-on-stdin/55701):
>
> Hi I am trying to use Julia for general purpose programming. I am trying to use Julia to create a faster version of a Perl script, which is supposed to create a report on a large logfile, piped into STDIN. The logfile does not fit into RAM, hence the data needs to be read from STDIN line-by-line from the data stream. My problem is that Julia (1.5.3) is twice as slow as Perl. I was able to identify one of the main problems, which is the code to read line-by-line from the input stream. I use …
