```
/var/log # ls -lh big.log
-rw-r--r-- 1 root root 1.5G Jan 1 00:06 big.log
/var/log # cat big.log | wc
10215006 119277417 1650294077
/var/log # time curl -s 'localhost:8080/log/read?file=big.log&mode=stream&limit=1000000000000' 2>/dev/null | wc
3852 19300 209632`
```
i've got a giant file (10215006 lines), and i'm trying to stream all of it and do a wc, and make sure it matches what i expect from `cat`. thinking theres a bug, or is it expected to short circuit early? (can we modify it so that i can have infinite matches?)
Also, the response of stream mode uses the same format as [SSE](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events#examples). So I don't think `wc` will give you the same numbers you got when you did `cat big.log | wc`.
thanks yichao! its a lot closer now, but seems to be off by a little bit.
two questions for you:
1. any theories why the number of lines is still slightly off?
2. any theories on why theres a large time difference between the two commands?
```
/var/log # time curl -s 'localhost:8080/log/read?file=big.log&mode=stream&limit=1000000000000000' | grep "data:" | wc
real 1m 13.69s
user 0m 4.77s
10174493 129306646 1810157535
sys 0m 2.14s
/var/log # time cat big.log | wc
10215008 119277419 1650294085
real 0m 7.03s
user 0m 0.18s
sys 0m 2.06s
/var/log #
```
heres also an interesting comparison:
```
/var/log # time curl -s 'localhost:8080/log/read?file=big.log&mode=stream&limit=1000000000000000&keyword=dax' | grep "data:"
data: {"line":"dax"}
data: {"line":"dax"}
data: {"hasMore":false}
real 0m 12.04s
user 0m 0.00s
sys 0m 0.01s
/var/log # time cat big.log | grep dax
dax
dax
real 0m 22.85s
user 0m 0.76s
sys 0m 3.97s
```
searching for a sparse keyword was faster than `cat`! and the `wc` will be expected (off by 1)
> any theories why the number of lines is still slightly off?
I saw the same issue when testing using the log file you provided in #1. There might be a bug somewhere, or the behavior of `wc` is different than expected in some cases. I will see if I can find where the root cause is.
> any theories on why theres a large time difference between the two commands?
There are mainly two factors that hurt performance the most:
1. Disk I/O.
2. Copy in memory.
My initial thought was that we were issuing multiple `read` calls (because we need to read the file from the back to the front) which might be slower than streaming the the whole file from beginning to end. But you also showed that it's faster when looking for keywords. So I guess the time spent in I/O might be similar.
Every time a line of log is found, the line will be converted from bytes to string and serialized to JSON before sending to response. These actions are likely to trigger unwanted copy of memory which slowed things down. If a keyword is given and a line does _not_ contain the keyword. No memory copy occurs. I think that's the reason why it's faster when searching for keyword.
Following up on
> any theories why the number of lines is still slightly off?
As shown in the below screenshot. I was comparing the output between my endpoint (left) and `tac` (right). Seems like an identical line was skipped by looking at the timestamp. I guess it happens each time we try to load remaining data into the buffer and somehow the offset is not correctly calculated.
<img width="1680" alt="image" src="https://github.com/user-attachments/assets/7487c42d-2ff6-46da-9b91-9788f1eb10a2" />
install.log looks good now!
@devMYC unfortunately, my big file is now stuck :D you can replicate the big.log via something like
``` tail install.log >> install.log``` and then control-c after a few seconds, repeat if needed.
```
/var/log # time curl -s 'localhost:8080/log/read?file=big.log&mode=stream&limit=1000000000000000' | grep "data:" | wc
^CCommand terminated by signal 2
real 4m 38.10s
user 0m 0.08s
sys 0m 0.11s
```
(gave up after 4 minutes, previously ran to completion in 1+ min)
edit --> actually i think its just really slow? maybe it does complete, will report back after i try again and wait longer.
Not sure if your log file is too big? I tried locally with a file size of 1.2GB and it terminates normally.
```
$ ls -ahl big.log
-rw-r--r-- 1 MYC staff 1.2G Jan 3 13:02 big.log
$ cat big.log | wc -l
8878076
$ time curl -s 'localhost:8080/log/read?file=big.log&mode=stream&limit=1000000000000000' | grep -E '^data:' | wc -l
8878077
real 1m40.769s
user 0m15.577s
sys 0m3.913s
```