Fix Tier 2 listing errors and easy Tier 3 items from future-concerns review

Fixes code listings that did not compile or contradicted their prose
across every chapter, regenerates gdb/ltrace/valgrind transcripts with
the real tools, corrects clear-cut technical errors, refreshes stale
material, cuts empty honors subsections and kernel TODOs, deletes the
orphaned honors/tcp.tex and introc/topics.tex, and applies the author's
answers to the questions raised. Trims resolved entries from
future-concerns-for-review.md and records figure bugs found while
assessing alt text.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Lawrence Angrave
2026-09-13 17:11:32 -05:00
co-authored by Claude Opus 5
parent 685a84d292
commit bd62265dbe
27 changed files with 324 additions and 897 deletions
+18 -17
View File
@@ -757,7 +757,7 @@ while (1) {
pthread_mutex_lock(&m);
count++;
pthread_cond_signal(&cv);
/* Even though the other thread is woken up it cannot not return */
/* Even though the other thread is woken up it cannot return */
/* from pthread_cond_wait until we have unlocked the mutex. This is */
/* a good thing! In fact, it is usually the best practice to call */
/* cond_signal or cond_broadcast before unlocking the mutex */
@@ -940,14 +940,15 @@ These examples are adapted from those.
Sequentially consistent is the simplest, least error-prone and most expensive model. This model says that any change that happens, all changes before it will be synchronized between all threads.
Suppose that \keyword{x} is atomic, \keyword{y} is an ordinary variable, and both start at 0.
\begin{verbatim}
Thread 1 Thread 2
1.0 atomic_store(x, 1)
1.1 y = 10 2.1 if (atomic_load(x) == 0)
1.2 atomic_store(x, 0); 2.2 y != 10 && abort();
1.1 y = 1; 2.1 if (atomic_load(x) == 2)
1.2 atomic_store(x, 2); 2.2 assert(y == 1);
\end{verbatim}
Will never quit.
The assert will never abort.
This is because either the store happens before the if statement in thread 2 and y == 1 or the store happens after and x does not equal 2.
\subsection{Relaxed}
@@ -964,11 +965,11 @@ One can have stale reads and writes, but after reading the new value, it won't b
\end{verbatim}
But that means that previous loads and stores don't need to affect other threads.
In the previous example, the code can now fail.
If the sequentially consistent example above used relaxed ordering, the code could now fail: thread 2 could see \keyword{x} == 2 but still read the old value \keyword{y} == 0, so the assert can abort.
\subsection{Acquire/Release}
The order of atomic variables doesn't need to be consistent -- meaning if atomic var y is assigned to 10 then atomic var x to be 0 those don't need to propagate, and a thread could get stale reads.
The order of atomic variables doesn't need to be consistent -- meaning if atomic var y is assigned to 1 then atomic var x to be 2 those don't need to propagate, and a thread could get stale reads.
Non-atomic variables have to get updated in all threads though.
\subsection{Consume}
@@ -1134,7 +1135,7 @@ Here are some concepts from queueing theory that you'll need to know that will h
A famous result in queueing theory is Little's Law which states $E[N] = \lambda E[W]$ meaning that the number of people waiting is the arrival rate times the expected waiting time (assuming the queue is in a steady state).
\item We won't make many assumptions about how much time it takes to run each process except that it will take a finite amount of time -- otherwise this gets almost impossible to evaluate.
We will denote two variables that $\frac{1}{\mu}$ is the mean of the waiting time and that the coefficient of variation $C$ is defined as $C^2 = \frac{var(S)}{E[S]^2}$ to help us control for processes that take a while to finish.
An important note is that when $C > 1$ we say that the running times of the process are variadic. We will note below that this rockets up the wait and response times for FCFS quadratically.
An important note is that when $C > 1$ we say that the running times of the process are highly variable. We will note below that this rockets up the wait and response times for FCFS quadratically.
\item $\rho = \frac{\lambda}{\mu} < 1$ Otherwise, our queue would become infinitely long
\item We will assume that there is one processor. This is known as an M/G/1 queue in queueing theory.
\item We'll leave the service time as an expectation $S$ otherwise we may run into over-simplifications with the algebra.
@@ -1161,7 +1162,7 @@ All results are from Jorma Virtamo's lectures on the matter \cite{virtamo}.
\]
The response time is simple to calculate, it is the expected number of people ahead of the process in the queue times the expected time to service each of those processes.
From Little's Law above, we can substitute that for this. Since we already know the value of the waiting time, we can reason about the response time as well.
\item A discussion of the results shows something cool discovered by Conway and Al \cite{conway1967theory}.
\item A discussion of the results shows something cool discovered by Conway et al. \cite{conway1967theory}.
Any scheduling discipline that isn't preemptive and doesn't take into account the run time of the process or a priority will have the same wait, response, and turnaround time.
We will often use this as a baseline.
\end{enumerate}
@@ -1382,7 +1383,7 @@ We will also introduce an additional term $C_i$ which denotes the variation amon
We incur the same cost on response time and then we have to suffer an additional cost based on what the probabilities are of lower priority jobs coming in and taking this job out.
That is what we call the average interruption time.
This follows the same laws as before.
Since we have a variadic, pyramid summation if we have a lot of jobs with small service times then the wait time goes down for both additive pieces.
Since we have a variable, pyramid summation if we have a lot of jobs with small service times then the wait time goes down for both additive pieces.
It can be analytically shown that this is better given certain probability distributions.
For example, try with the uniform versus FCFS or the non preemptive version.
What happens?
@@ -1411,8 +1412,8 @@ Datagrams are formatted as such
\end{figure}
\begin{enumerate}
\item The first octet is the version number, either 4 or 6
\item The next octet is how long the header is.
\item The first four bits are the version number, which is 4 for IPv4.
\item The next four bits, completing the first octet, are the header length (IHL), counted in 4-octet words.
Although it may seem that the header is a constant size, you can include optional parameters to augment the path that is taken or other instructions.
\item The next two octets specify the total length of the datagram.
This means this is the header, the data, the footer, and the padding.
@@ -1420,9 +1421,9 @@ Datagrams are formatted as such
\item The next two are the Identification number.
IP handles taking packets that are too big to be sent over the physical wire and chunks them up.
As such, this number identifies what datagram this originally belonged to.
\item The next octet is various bit flags that can be set.
\item The next octet and half is fragment number.
If this packet was fragmented, this is the number this fragment represents.
\item The next three bits are flags that can be set, such as ``don't fragment'' and ``more fragments''.
\item The next thirteen bits, completing those two octets, are the fragment offset.
If this packet was fragmented, this says where this fragment's data belongs in the original datagram, counted in units of 8 octets.
\item The next octet is time to live.
So this is the number of "hops" (travels over a wire) a packet is allowed to go.
This is set because different routing protocols could cause packets to go in circles, the packets must be dropped at some point.
@@ -1439,7 +1440,7 @@ Datagrams are formatted as such
There is no verification of this, so one host can pretend to be any IP address possible.
\item The destination address is where you want the packet to be sent to.
Destinations are crucial to the routing process.
\item Additional options: Hosts of additional options, this is variadic in size.
\item Additional options: Hosts of additional options, this is variable in size.
\item Footer: A bit of padding to make sure your data is a multiple of 4 octets.
\item After: Your data! All data of higher-order protocols are put following the header.
\end{enumerate}
@@ -1507,7 +1508,7 @@ Note that this will cause undue stress on the network, so a series of multicasts
When it comes to Event-Driven IO, the name of the game is to be fast.
One extra system call is considered slow.
OpenBSD and FreeBSD have an arguably better model of asynchronous IO from the kqueue model.
Kqueue is a system call that is exclusive to the BSDs and MacOs.
Kqueue is a system call that is exclusive to the BSDs and macOS.
It allows you to modify file descriptor events and read file descriptors all in a single call under a unified interface.
So what are the benefits?
+47 -38
View File
@@ -11,7 +11,7 @@ This section is a short review of System Architecture topics that you'll need fo
What is assembly? Assembly is the lowest that you'll get to machine language without writing 1's and 0's.
Each computer has an architecture, and that architecture has an associated assembly language.
Each assembly command has a 1:1 mapping to a set of 1's and 0's that tell the computer exactly what to do.
For example, the following in the widely used x86 Assembly language adds one to the memory address 20 \cite{wiki:xxx} -- you can also look in \cite{guide2011intel} Section 2A under the add instruction though it is more verbose.
For example, the following in the widely used x86 Assembly language adds one to the byte at memory address \keyword{0x20} \cite{wiki:xxx} -- you can also look in \cite{guide2011intel} Section 2A under the add instruction though it is more verbose.
\begin{lstlisting}[language={[x86masm]Assembler}]
add BYTE PTR [0x20], 1
@@ -22,7 +22,7 @@ Serious implications arise for race conditions and atomic operations.
\subsection{Atomic Operations}
An operation is atomic if no other processor should interrupt it. Take for example the above assembly code to add one to a register.
An operation is atomic if no other processor should interrupt it. Take for example the above assembly code to add one to a memory address.
In the architecture, it may actually have a few different steps on the circuit.
The operation may start by fetching the value of the memory from the stick of ram, then storing it in the cache or a register, and then finally writing back \cite{schweizer2015evaluating} -- under the description for \textit{fetch-and-add} though your micro-architecture may vary.
Or depending on performance operations, it may keep that value in cache or in a register which is local to that process -- try dumping the \keyword{-O2} optimized assembly of incrementing a variable.
@@ -79,7 +79,8 @@ Operating system developers and instruction set developers alike didn't like the
What safely means is obviously out of the scope for this class, but it persists.
\subsection{Optional: Hyperthreading}
Hyperthreading is a new technology and is in no way shape or form multithreading.
Hyperthreading is Intel's name for simultaneous multithreading (SMT), a hardware technology that Intel first shipped in 2002.
Despite the name, it is not the same thing as software multithreading: one physical core presents two logical CPUs that share its execution units.
Hyperthreading allows one physical core to appear as many virtual cores to the operating system \cite[P.51]{guide2011intel}.
The operating system can then schedule processes on these virtual cores and one core will execute them.
Each core interleaves processes or threads.
@@ -123,7 +124,7 @@ If you don't want to type your password out every time, you can generate an ssh
If you still think that that is too much typing, you can always alias hosts.
You may need to restart your VM or reload sshd for this to take effect.
The config file is available on Linux and Mac distros.
For Windows, you'll have to use the Windows Linux Subsystem or configure any aliases in PuTTY
For Windows, you'll have to use the Windows Subsystem for Linux (WSL) or configure any aliases in PuTTY
\begin{lstlisting}[language=bash]
> cat ~/.ssh/config
@@ -136,9 +137,10 @@ Host vm
\subsection{git}
What is `git`? Git is a version control system. What that means is git stores the entire history of a directory. We refer to the directory as a repository. So what you need to know is a few things. First, create your repository with the repo creator. If you haven't already signed into enterprise GitHub, make sure to do so otherwise your repository won't be created for you. After that, your repository is created on the server. Git is a decentralized version control system, meaning that you'll need to get a repository onto your VM. We can do this with a clone. Whatever you do, \textbf{do not go through the README.md tutorial}.
Replace \keyword{<semester>} with the current semester code (for example, \keyword{fa26}) and \keyword{<netid>} with your NetID.
\begin{lstlisting}[language=bash]
$ git clone https://github.com/illinois-cs-coursework/fa23_cs341_<netid>
$ git clone https://github.com/illinois-cs-coursework/<semester>_cs341_<netid>
\end{lstlisting}
This will create a local repository. The workflow is you make a change on your local repository, add the changes to a current commit, actually commit, and push the changes to the server.
@@ -200,7 +202,7 @@ nothing to commit, working directory clean
Don't panic, but your repository may be in an unworkable state.
If you aren't nearing a deadline, come to office hours or ask your question on Edstem, and we'd be happy to help.
In an emergency scenario, delete your repository and re-clone (you'll have to add the \keyword{release} as above).
In an emergency scenario, delete your repository and re-clone.
\textbf{This will lose any local uncommitted changes. Make sure to copy any files you were working on outside the directory, remove and copy them back in.}
If you want to learn more about git, there are all but an endless number of tutorials and resources online that can help you.
@@ -335,7 +337,7 @@ Suppose we have a simple program like this:
void dummy_function() {
int* x = malloc(10 * sizeof(int));
x[10] = 0; // error 1: Out of bounds write, as you can see here we write to an out of bound memory address.
} // error 2: Memory Leak, x is allocated at function exit.
} // error 2: Memory Leak, x is not freed at function exit.
int main(void) {
dummy_function();
@@ -348,33 +350,37 @@ Let's see what Valgrind will output.
\begin{lstlisting}[language=C]
==29515== Memcheck, a memory error detector
==29515== Copyright (C) 2002-2015, and GNU GPL'd, by Julian Seward et al.
==29515== Using Valgrind-3.11.0 and LibVEX; rerun with -h for copyright info
==29515== Copyright (C) 2002-2022, and GNU GPL'd, by Julian Seward et al.
==29515== Using Valgrind-3.22.0 and LibVEX; rerun with -h for copyright info
==29515== Command: ./a
==29515==
==29515== Invalid write of size 4
==29515== at 0x400544: dummy_function (in /home/rafi/projects/exocpp/a)
==29515== by 0x40055A: main (in /home/rafi/projects/exocpp/a)
==29515== Address 0x5203068 is 0 bytes after a block of size 40 alloc'd
==29515== at 0x4C2DB8F: malloc (in /usr/lib/valgrind/vgpreload_memcheck-amd64-linux.so)
==29515== by 0x400537: dummy_function (in /home/rafi/projects/exocpp/a)
==29515== by 0x40055A: main (in /home/rafi/projects/exocpp/a)
==29515== at 0x108774: dummy_function (in /home/user/a)
==29515== by 0x10878F: main (in /home/user/a)
==29515== Address 0x4a7e068 is 0 bytes after a block of size 40 alloc'd
==29515== at 0x4885250: malloc (in /usr/libexec/valgrind/vgpreload_memcheck-arm64-linux.so)
==29515== by 0x108767: dummy_function (in /home/user/a)
==29515== by 0x10878F: main (in /home/user/a)
==29515==
==29515==
==29515== HEAP SUMMARY:
==29515== in use at exit: 40 bytes in 1 blocks
==29515== total heap usage: 1 allocs, 0 frees, 40 bytes allocated
==29515==
==29515== 40 bytes in 1 blocks are definitely lost in loss record 1 of 1
==29515== at 0x4885250: malloc (in /usr/libexec/valgrind/vgpreload_memcheck-arm64-linux.so)
==29515== by 0x108767: dummy_function (in /home/user/a)
==29515== by 0x10878F: main (in /home/user/a)
==29515==
==29515== LEAK SUMMARY:
==29515== definitely lost: 40 bytes in 1 blocks
==29515== indirectly lost: 0 bytes in 0 blocks
==29515== possibly lost: 0 bytes in 0 blocks
==29515== still reachable: 0 bytes in 0 blocks
==29515== suppressed: 0 bytes in 0 blocks
==29515== Rerun with --leak-check=full to see details of leaked memory
==29515==
==29515== For counts of detected and suppressed errors, rerun with: -v
==29515== ERROR SUMMARY: 1 errors from 1 contexts (suppressed: 0 from 0)
==29515== For lists of detected and suppressed errors, rerun with: -s
==29515== ERROR SUMMARY: 2 errors from 2 contexts (suppressed: 0 from 0)
\end{lstlisting}
\textbf{Invalid write}: It detected our heap block overrun, writing outside of an allocated block.
@@ -460,13 +466,13 @@ $ gdb --args ./main
(gdb) r
[...]
Program received signal SIGTRAP, Trace/breakpoint trap.
main () at main.c:6
6 val = 7;
main () at main.c:5
5 val = 7;
(gdb) p val
$1 = 42
\end{lstlisting}
You can also set breakpoints programmatically.
You can also set breakpoints interactively from within gdb.
Assume that we have no optimization and the line numbers are as follows
\begin{lstlisting}[language=bash]
@@ -484,6 +490,8 @@ $ gcc main.c -g -o main
$ gdb --args ./main
(gdb) break main.c:4
[...]
(gdb) r
[...]
(gdb) p val
$1 = 42
\end{lstlisting}
@@ -530,7 +538,7 @@ Breakpoint 1, main () at main.c:4
(gdb)
\end{lstlisting}
Here, by using the \keyword{x} command with parameters \keyword{16xb}, we can see that starting at memory address \keyword{0x7fff5fbff9c} (value of \keyword{bad\_string}), \keyword{printf} would actually see the following sequence of bytes as a string because we provided a malformed string without a null terminator.
Here, by using the \keyword{x} command with parameters \keyword{16xb}, we can see that starting at memory address \keyword{0x7fff5fbff9cd} (value of \keyword{bad\_string}), \keyword{printf} would actually see the following sequence of bytes as a string because we provided a malformed string without a null terminator.
\subsection{Involved gdb example}
@@ -610,11 +618,11 @@ Okay, flip the sign it should work now right?
(gdb) print/x deg # print the hex value of degree
$1 = 0x167
(gdb) print (31415/1000)
$2 = 0x31
$2 = 31
(gdb) print (31415/1000.0)
$3 = 201.749
$3 = 31.414999999999999
(gdb) print (31415.0/10000.0)
$4 = 3.1414999999999999
$4 = 3.1415000000000002
\end{lstlisting}
@@ -678,14 +686,15 @@ int main() {
\begin{lstlisting}[language=bash]
> ltrace ./a.out
__libc_start_main(0x8048454, 1, 0xbfc19db4, 0x80484c0, 0x8048530 <unfinished ...>
fopen("I don't exist", "r") = 0x0
fwrite("Invalid Write\n", 1, 14, 0x0 <unfinished ...>
--- SIGSEGV (Segmentation fault) ---
+++ killed by SIGSEGV +++
__libc_start_main(0xaaaaab1a0818, 1, 0xffffc0fdaa28, 0 <unfinished ...>
fopen("I don't exist", "r") = 0
fputc('a', 0 <no return ...>
--- SIGSEGV (Segmentation fault) ---
+++ killed by SIGSEGV +++
\end{lstlisting}
Notice that the compiler has replaced the one-character \keyword{fprintf} with a call to \keyword{fputc}, and that the \keyword{FILE*} it is given is \keyword{0} (\keyword{NULL}), the value \keyword{fopen} returned.
ltrace output can clue you in to weird things your program is doing live.
Unfortunately, ltrace can’t be used to inject faults, meaning that ltrace can tell you what is happening, but it can't tamper with what is already happening.
@@ -724,10 +733,10 @@ read(3, "# C Datastructures\n\n[![Build Sta"..., 8192) = 1250
\end{lstlisting}
Newer versions of strace can actually inject faults into your program.
Newer versions of strace can actually inject faults into your program with the \keyword{-e inject=} option.
For example, \keyword{strace -e inject=read:error=EIO head README.md} makes every \keyword{read} call fail with \keyword{EIO}.
This is useful when you want to occasionally make reads and writes fail for example in a networking application, which your program should handle.
The problem is as of early 2019, that version is missing from Ubuntu repositories.
Meaning that you'll have to install it from the source.
The strace in current Ubuntu releases supports fault injection, so the package installed above is all you need.
\subsection{printfs}
@@ -806,7 +815,7 @@ if(!fork()) {
}
if(!fork()) {
execlp("make","make", "snowman", (char*)0); execlp("make","make", (char*)0));
execlp("make","make", "snowman", (char*)0); execlp("make","make", (char*)0);
}
exit(0);
@@ -818,7 +827,7 @@ exit(0);
int main(int argc, char** argv) {
puts("Great! We have plenty of useful resources for you, but it's up to you to");
puts(" be an active learner and learn how to solve problems and debug code.");
puts("Bring your near-completed answers the problems below");
puts("Bring your near-completed answers to the problems below");
puts(" to the first lab to show that you've been working on this.");
printf("A few \"don't knows\" or \"unsure\" is fine for lab 1.\n");
puts("Warning: you and your peers will work hard in this class.");
@@ -893,11 +902,11 @@ char *ptr = "hello";
\item What is the value of the variable \keyword{str\_size}?
\begin{lstlisting}[language=C]
ssize_t str_size = sizeof("Hello\0World")
ssize_t str_size = sizeof("Hello\0World");
\end{lstlisting}
\item What is the value of the variable \keyword{str\_len}
\begin{lstlisting}[language=C]
ssize_t str_len = strlen("Hello\0World")
ssize_t str_len = strlen("Hello\0World");
\end{lstlisting}
\item Give an example of X such that \keyword{sizeof(X)} is 3.
\item Give an example of Y such that \keyword{sizeof(Y)} might be 4 or 8 depending on the machine.
@@ -1010,7 +1019,7 @@ Oh, and did I mention that this is an easy way to score points with your interns
\item Did I Google the error message and a few permutations thereof if necessary? How about StackOverflow.
\item Did I try commenting out, printing, and/or stepping through parts of the code bit by bit to find out precisely where the error occurs?
\item \textbf{Did I commit my code to git in case the TAs need more context?}
\item Did I include the console/GDB/Valgrind output **AND** code surrounding the bug in my class forum post?
\item Did I include the console/GDB/Valgrind output \textbf{AND} code surrounding the bug in my class forum post?
\item Have I fixed other segmentation faults unrelated to the issue I'm having?
\item Am I following good programming practice? (i.e. encapsulation, functions to limit repetition, etc)
\end{enumerate}
+4 -4
View File
@@ -34,7 +34,7 @@ It is a simple yet powerful tool to illustrate how interacting processes can dea
If a process is \emph{using} a resource, an arrow is drawn from the resource node to the process node.
If a process is \emph{requesting} a resource, an arrow is drawn from the process node to the resource node.
If there is a cycle in the Resource Allocation Graph and each resource in the cycle provides only one instance, then the processes will deadlock.
For example, if process 1 holds resource A, process 2 holds resource B and process 1 is waiting for B and process 2 is waiting for A, then processes 1 and 2 will be deadlocked \ref{ragfigure}.
For example, if process 1 holds resource A, process 2 holds resource B and process 1 is waiting for B and process 2 is waiting for A, then processes 1 and 2 will be deadlocked (see Figure~\ref{ragfigure}).
We'll make the distinction that the system is in deadlock by definition if all workers cannot perform an operation other than waiting.
So, as long as every resource offers a single instance, detecting deadlock means searching the graph for a cycle.
Both the processes and the resources are nodes, and every edge has a direction, so this is cycle detection in a \emph{directed} graph.
@@ -171,7 +171,7 @@ Consider the scenario where two students need to write both pen and paper and th
Breaking mutual exclusion means that the students share the pen and paper.
Breaking circular wait could be that the students agree to grab the pen then the paper.
As proof by contradiction, say that deadlock occurs under the rule and the conditions.
Without loss of generality, that means a student would have to be waiting on a pen while holding the paper and the other waiting on a pen and holding the paper.
Without loss of generality, that means a student would have to be waiting on the pen while holding the paper and the other waiting on the paper while holding the pen.
We have contradicted ourselves because one student grabbed the paper without grabbing the pen, so deadlock fails to occur.
Breaking hold and wait could be that the students try to get the pen and then the paper and if a student fails to grab the paper then they release the pen.
This introduces a new problem called \textit{livelock} which will be discussed later.
@@ -382,7 +382,7 @@ The system is about to deadlock, but the approach resolves it.
\begin{figure}[H]
\centering
\includegraphics[width=.9\textwidth]{deadlock/drawings/dining_stalling.eps}
\caption{Stallings solution almost deadlock}
\caption{Stallings' solution almost deadlock}
\end{figure}
@@ -424,7 +424,7 @@ This prioritizes philosophers that have already eaten but can be made fairer by
\begin{figure}[H]
\centering
\includegraphics[width=.9\textwidth]{deadlock/drawings/dining_partial.eps}
\caption{Stallings solution partial deadlock}
\caption{Dijkstra's partial ordering solution}
\end{figure}
There are a few other solutions (clean/dirty forks and the actor model) in the appendix.
+12 -12
View File
@@ -367,9 +367,9 @@ closedir(dirp);
\end{lstlisting}
One final note of caution.
\keyword{readdir} is not thread-safe!
You shouldn't use the re-entrant version of the function.
Synchronizing the filesystem within a process is important, so use locks around \keyword{readdir}.
\keyword{readdir} is not guaranteed to be thread-safe when several threads share the same directory stream.
Don't reach for the re-entrant version \keyword{readdir\_r}; it is deprecated in glibc.
Instead, threads that read from separate \keyword{DIR*} streams need no extra care, while threads that share one \keyword{DIR*} stream must use a lock around each call to \keyword{readdir}.
See the \href{https://linux.die.net/man/3/readdir}{man page of readdir} for more details.
@@ -423,7 +423,7 @@ link("file1.txt", "blip.txt");
\begin{verbatim}
$ ln -s file1.txt file2.txt
$ ls -i file1.txt blip.txt
$ ls -i file1.txt file2.txt blip.txt
134235 file1.txt
134236 file2.txt
134235 blip.txt
@@ -438,13 +438,13 @@ $ cat file1.txt
file1!
edited file2
$ cat file2.txt
I'm file1!
file1!
edited file2
$ cat blip.txt
file1!
edited file2
$ readlink myfile.txt
file2.txt
$ readlink file2.txt
file1.txt
\end{verbatim}
Note that \keyword{file2.txt} and \keyword{file1.txt} have different inode numbers, unlike the hard link, \keyword{blip.txt}.
@@ -690,7 +690,7 @@ The octal number is the sum of three values given to the three types of permissi
Example: \keyword{chmod 755 myfile}
\begin{enumerate}
\item r + w + x = digit * user has 4+2+1, full permission
\item user has 4+2+1, read, write and execute (full) permission
\item group has 4+0+1, read and execute permission
\item all users have 4+0+1, read and execute permission
\end{enumerate}
@@ -808,7 +808,7 @@ Device & Use Case \\ \hline
If we want a continuous stream of 0s, we can run \keyword{cat /dev/zero}.
Another example is the file \keyword{/dev/null}, a great place to store bits that you never need to read.
Bytes sent to \keyword{/dev/null/} are never stored and simply discarded.
Bytes sent to \keyword{/dev/null} are never stored and simply discarded.
A common use of \keyword{/dev/null} is to discard standard output.
For example,
@@ -947,7 +947,7 @@ This means the speed of the transfer is unaffected by hardware power.
\begin{lstlisting}[language=bash]
$ mkdir arch
$ sudo mount -o loop archlinux-2015.04.01-dual.iso ./arch
$ sudo mount -o loop archlinux-2014.11.01-dual.iso ./arch
$ cd arch
\end{lstlisting}
@@ -1161,7 +1161,7 @@ It is laid out sequentially on disk, and the first section is the superblock.
The superblock stores important metadata about the entire filesystem.
Since we want to be able to read this block before we know anything else about the data on disk, this needs to be in a well-known location so the start of the disk is a good choice.
After the superblock, we'll keep a map of which inodes are being used.
The nth bit is set if the nth inode -- $0$ being the inode root -- is being used.
The nth bit is set if the nth inode -- $0$ being the root inode -- is being used.
Similarly, we store a map recording which data blocks are used.
Finally, we have an array of inodes followed by the rest of the disk - implicitly partitioned into data blocks.
One data block may be identical to the next from the perspective of the hardware components of the disk.
@@ -1257,7 +1257,7 @@ We perform our write and go on our merry way.
Some questions to consider.
\begin {itemize}
\begin{itemize}
\item How would a program perform a write that goes across data block boundaries?
\item How would a program perform a write after adding the offset would extend the length of the file?
\item How would a program perform a write where the offset is greater than the length of the original file?
+31 -595
View File
@@ -16,29 +16,45 @@ file is the leftovers.
- **These findings are unverified unless marked otherwise.** They were
produced by an automated pass and will contain false positives. Confirm
before acting, especially on technical claims.
- The section immediately below is the exception: those items were checked
by hand against the source and the published PDF.
- Items that have since been fixed (PR #235 and the Tier 2 PR that
followed it) have been removed, so line numbers may have drifted.
Locate items by their quoted text.
---
## Verified by hand
## Figures
### Two source files are orphaned — written, but never built
Found while viewing every figure to assess alt text (none of the 48
figures has any). Fix these before writing alt text that states the
figures' contents.
- `introc/topics.tex` is never `\input`, and its content is duplicated
verbatim inside `introc/introc.tex`. Editing the standalone file has no
effect on the book.
- `honors/tcp.tex` is never `\input` by `honors/honors.tex`, and is a single
truncated sentence.
### introc/c_memory_model.tex:66-98 — the three memory-model figures are blank
`memory_model_empty.eps`, `memory_model_length.eps` and
`memory_model_full.eps` are identical apart from a timestamp: 11 empty
boxes, with no "0006" and no string, contradicting their captions. Commit
a620438 (2020-01-12) replaced them. The 521ba93 originals had the content,
but spelled the string "bhuvan" rather than "person", so they need
redrawing.
Decide whether each should be wired in or deleted; leaving unreferenced
sources invites edits that silently do nothing.
### scheduling/drawings/psjf.eps — P5 gets 4 s of CPU
The text gives P5 5000ms, but its bars total 4 s.
### `honors/containers.tex` has three empty subsections
### ipc/drawings/three_address_split.eps — index labels look swapped
The "Index 1" / "Index 2" labels appear reversed relative to the prose.
"Linux Namespaces", "Building a container from scratch" and "Containers in
the wild" are headings with no body text. They render as empty sections in
the published book.
### ipc/drawings/frame_table.eps — numbers differ from the prose
The figure maps page 1 → frame 30 and page 2 → 24; the prose says 1 → 45
and 2 → 30.
### Alt text — mechanism
`\includegraphics[alt={...}]` compiles on the CI toolchain (TeX Live 2023)
but is currently discarded everywhere: the PDF is untagged, and pandoc 2.7
(EPUB/wiki) uses the caption as alt and drops `alt=`. PDF tagging
(`\DocumentMetadata`) fails with the book's listings setup under TL2023.
Pandoc 3.x does honour `alt=`, but then figures without it get empty alt,
which the EPUB filters' `NoAltTagException` rejects. Draft alt text for 45
of the 48 figures exists from the assessment and can be applied when
wanted.
---
@@ -78,99 +94,10 @@ The Authors section renders a markdown file in a monospaced listing environment.
## background
### Stale per-student repository URL — background/background.tex:141
```
$ git clone https://github.com/illinois-cs-coursework/fa23_cs341_<netid>
```
Hardcodes the FA23 semester prefix. For FA26 this should be `fa26_cs341_<netid>` (or be written generically). Flagged as requested; needs a human to confirm the semester naming scheme actually in use.
### "Hyperthreading is a new technology" — background/background.tex:82
Hyper-Threading shipped commercially in 2002. Calling it "new" is dated. Also the subsection asserts it "is in no way shape or form multithreading", which is a strong claim a student may find confusing given the name. Needs an author decision on rewording.
### gdb transcript values look wrong — background/background.tex:611-617
```
(gdb) print (31415/1000)
$2 = 0x31
(gdb) print (31415/1000.0)
$3 = 201.749
```
`31415/1000` is 31 (`0x1f`, not `0x31`), and gdb's `print` does not inherit the `/x` format from the previous command anyway. `31415/1000.0` is `31.415`, not `201.749`. These appear to be fabricated/garbled outputs; since this is the punchline of the debugging walkthrough, a human should regenerate the real transcript. (Left untouched: inside `lstlisting`.)
### Valgrind example comment states the wrong reason — background/background.tex:338
```
} // error 2: Memory Leak, x is allocated at function exit.
```
The leak is that `x` is *not freed* at function exit; "is allocated" doesn't describe the error. Inside a listing so left alone, but it teaches the wrong wording.
### ltrace example output does not match its source — background/background.tex:670-687
The C snippet calls `fprintf(fp, "a")`, but the ltrace output shows `fwrite("Invalid Write\n", 1, 14, 0x0 ...)`. The string and length do not correspond to the program shown. A human should regenerate or reconcile the example.
### "Setting breakpoints programmatically" paragraph is self-contradictory — background/background.tex:443-489
The paragraph titled "Setting breakpoints programmatically" first shows the `asm("int $3")` technique (which *is* the programmatic one), then says "You can also set breakpoints programmatically" and demonstrates `break main.c:4` — an interactive, non-programmatic breakpoint. The labels appear swapped. Confusing for a first-time gdb user.
### Truncated memory address in prose — background/background.tex:533
> "starting at memory address \keyword{0x7fff5fbff9c}"
The gdb output above it shows `0x7fff5fbff9cd` (14 hex digits vs 13). Looks like a dropped character, but it is an address inside a `\keyword{}` so I did not touch it.
### "add one to the memory address 20" — background/background.tex:14
The instruction is `add BYTE PTR [0x20], 1`, i.e. address `0x20` = 32 decimal. Prose says "memory address 20", which a student will read as decimal 20. Suggest "0x20".
### Broken logic in the git-status troubleshooting flow — background/background.tex:168-201
"If you are currently on a branch, and you don't see either \<A\> or \<B\>" ... then line 193 continues "And something like \<C\>". The condition never resolves grammatically or logically: is the trigger *not* seeing A/B, or *seeing* C? As written a student can't tell what state means "don't panic, but your repository may be in an unworkable state". Needs an author rewrite.
### Undefined reference to "the \keyword{release}" — background/background.tex:203
> "delete your repository and re-clone (you'll have to add the \keyword{release} as above)"
Nothing "above" explains a `release` remote or branch — the git subsection only covers clone/add/commit/push and mentions a feedback branch. Missing context a student following this emergency procedure would need.
### Outdated strace claim — background/background.tex:729
> "The problem is as of early 2019, that version is missing from Ubuntu repositories."
Fault injection (`-e inject=`) has been in Ubuntu's strace for many releases now. Time-stamped claim that is almost certainly no longer true.
### Windows Subsystem for Linux named incorrectly — background/background.tex:126
> "you'll have to use the Windows Linux Subsystem"
The product is "Windows Subsystem for Linux" (WSL). Left as-is because it is a proper-noun/technical name rather than a plain typo.
### Markdown emphasis leaking into LaTeX — background/background.tex:1013
> "Did I include the console/GDB/Valgrind output **AND** code surrounding the bug"
`**AND**` is Markdown syntax; in LaTeX it renders as literal asterisks. Presumably should be `\textbf{AND}`. Not changed because it alters markup, not prose.
### Extra closing parenthesis in Homework 0 listing — background/background.tex:809
```
execlp("make","make", (char*)0));
```
One `)` too many. Might be a deliberate "spot the bug" element of the lyrics puzzle, so I left it — but it is worth confirming it is intentional.
### Grammatical error inside the HW0 `minted` block — background/background.tex:821
```
puts("Bring your near-completed answers the problems below");
```
Missing "to" ("answers to the problems below"). This is student-facing output text, but it lives inside a code listing, so I did not edit it.
### Generic Edstem link — background/background.tex:847-848
> "Use the current semester's CS341 Edstem: \url{https://edstem.org/}"
@@ -197,22 +124,6 @@ The clause "and is a privileged operation" has no valid subject (the kernel is n
"...so the program compiles on missing variable because the program will reference a variable in the system or another file."
"compiles on missing variable" is not grammatical and the causal clause is circular. Needs a real rewrite of the explanation.
### introc/language_facilities.tex:240 — mislabelled code comment
The fourth example in the if/else listing is commented `// (1)` but should be `// (4)` (it matches "an if with an else if and else" from line 214). It is inside a listing, so I did not touch it, but it looks like a genuine copy-paste error rather than a deliberate bug example.
### introc/language_facilities.tex:250-252 — `inline` description contradicts itself / wrong term
"tells the compiler it's okay to omit the C function call procedure and \"paste\" the code in the callee. Instead, the compiler is hinted at substituting..."
The code is pasted into the *caller*, not the callee. Also "Instead," makes no sense between the two sentences since they say the same thing. Technical claim — needs an author.
### introc/language_facilities.tex:255-256 — `max` returns the minimum
"inline int max(int a, int b) { return a < b ? a : b; }" returns the smaller value. This is in a listing so I left it, but nothing in the surrounding text suggests the bug is intentional (the example is about `inline`, not about a bug).
### introc/language_facilities.tex:265 — `restrict` wording
"tells the compiler that this particular memory region shouldn't overlap with all other memory regions" — "with all other" should probably be "with any other"; as written it is ambiguous and technically weaker than intended. Left alone as a semantics claim.
### introc/language_facilities.tex:347-352 — count mismatch
"\keyword{static} is a type specifier with three meanings." but only two enumerated items follow. Either a third meaning was dropped or the count is wrong.
### introc/language_facilities.tex:369 — ungrammatical struct definition
"C-structs are contiguous regions of memory that one can access specific elements of each memory as if they were separate variables."
The relative clause is broken; needs rewriting by someone who knows the intended sentence.
@@ -225,33 +136,14 @@ The relative clause is broken; needs rewriting by someone who knows the intended
"know that most functions in C handle errors return oriented."
Probably "handle errors in a return-oriented way". Needs an author's wording.
### introc/common_c_functions.tex:198-199 — "Instead of" appears to be backwards
"Also naturally like \keyword{printf}, \keyword{scanf} functions require valid pointers. Instead of pointing to valid memory, they need to also be writable."
The intended meaning is surely "In addition to pointing to valid memory, they need to be writable."
### introc/common_c_functions.tex:211 — prose refers to a variable that isn't in the example
"We wanted to write the character value into c..." but the listing declares `char type`, not `c`.
### introc/common_c_functions.tex:317 — broken sentence
"The caller has to be careful from a valid 0 and an error."
Presumably "has to distinguish a valid 0 from an error."
### introc/common_c_functions.tex:325 — missing semicolon in listing
"errno = 0" in the strtol errno-trampoline example has no semicolon, so the snippet would not compile. In a listing, and this example is not presented as buggy code, so it looks accidental.
### introc/common_c_functions.tex:240 — strlen/strcmp prototypes
"\keyword{int strlen(const char *s)}" — the real prototype returns `size_t`. Since these are presented as reference prototypes rather than bug examples, a human may want to correct it. Same section, line 242: strcmp is documented as returning exactly -1/0/1, which is not guaranteed by the standard (only the sign is).
### introc/common_c_functions.tex:341 — dangling fragment
"\keyword{memcpy} and \keyword{memmove} both in \keyword{string.h}?"
This is not a sentence and the itemize ends on it. Possibly a leftover note ("Why are memcpy and memmove both in string.h?").
### introc/c_memory_model.tex:9 — dangling reference
"Consider the contact struct declared above." Nothing is declared above — the struct listing comes *after* this sentence, and this is the first mention of it in the chapter.
### introc/c_memory_model.tex:75-76 — endianness claim vs. the figure
"We will assume that our machine is big endian. This means that the least significant byte is last." The following figure caption says the four bytes are "filled with 0006", which is the big-endian layout, so the text and figure agree — but the memory-layout example at lines 35-38 and the "zero length array" hack do not depend on endianness at all. Worth a human check that this aside is not confusing students.
### introc/c_memory_model.tex:66-98 — figures have no alt text
The three \includegraphics figures (memory_model_empty.eps, memory_model_length.eps, memory_model_full.eps) rely on captions only. The captions are descriptive, but there is no alt-text mechanism for screen readers.
@@ -259,29 +151,6 @@ The three \includegraphics figures (memory_model_empty.eps, memory_model_length.
"In addition to adding to an integer, pointers can be added to."
Presumably "In addition to being able to add integers to integers, you can add an integer to a pointer." As written it is close to meaningless.
### introc/pointers.tex:115 — Markdown markup left in LaTeX
"we are increasing the **integer** pointer by 1" uses Markdown bold inside a .tex file; it will render literally as asterisks. Should probably be \textbf{integer}.
### introc/pointers.tex:116 — attribution of the void-pointer rule
"POSIX standards forbid arithmetic on void pointers." It is ISO C, not POSIX, that makes arithmetic on `void*` a constraint violation (POSIX actually leans the other way for some interfaces). Technical claim — left alone.
### introc/logic_and_program_flow_mistakes.tex:26 — confusing and possibly wrong claim
"Most modern compilers disallows assigning variables a condition without parenthesis."
Compilers *warn* (-Wparentheses), they do not disallow it; and "assigning variables a condition" is garbled. I fixed only the verb agreement.
### introc/logic_and_program_flow_mistakes.tex:52 — logic inverted
"The compiler fails to catch this error because the programmer omitted the valid function prototype by including \keyword{time.h}."
The prototype is missing because the programmer *did not* include time.h. As written it says the opposite.
### introc/logic_and_program_flow_mistakes.tex:29 — example does not compile
"if (42 = answer)" is offered as "the quick way to fix" the previous bug, but assigning to a constant is a compile error, which is the actual point being made. The surrounding text never says so explicitly, so a student may read this as recommended working code.
### introc/introc.tex:25-73 vs introc/topics.tex — duplicated content
The "Topics" itemize in introc.tex:25-73 is byte-for-byte the same list as introc/topics.tex, and topics.tex is never \input by introc.tex. One of the two is dead/duplicated.
### introc/history_of_c.tex:6 — dated hardware claim
"it was made to target the most popular computers at the time, such as the PDP-7." C was developed on the PDP-11; the PDP-7 was the machine for the earlier B/assembly UNIX. Historical claim worth checking against the cited source.
### introc/crash_course_introduction_to_c.tex:29 — flushing claim
"If the newline isn't included, the buffer will not be flushed (i.e. the write will not complete immediately)." True only for a line-buffered stdout, and the buffer is still flushed at exit. common_c_functions.tex:99-101 states the nuanced version; this simplified claim may mislead.
@@ -292,17 +161,6 @@ The "Topics" itemize in introc.tex:25-73 is byte-for-byte the same list as intro
## processes
### processes/processes.tex:483,494 — example code calls `fork` without parentheses
Both fork-and-FILEs snippets use `if(!fork) {`, which takes the address of the function (always
non-NULL, so the branch never runs) instead of calling it. Presumably should be `if (!fork())`. It is
inside a listing, so I left it, but it is very likely an unintended bug rather than a teaching bug.
### processes/processes.tex:520,550 — unbalanced parentheses in the getline loop
`while((nread = getline(&buffer, &buffer_cap, file) != -1) {` has three opening parens and two
closing; it will not compile. Also the `!= -1` binds to the `getline` result before the assignment,
so `nread` gets 0/1. The surrounding prose is about buffering/fork semantics, so this looks accidental
rather than the intended lesson.
### processes/processes.tex:225 — sentence appears to contradict/duplicate line 223
"Another is to use the built-in \keyword{exec} command to kill all the user processes (you only have
one attempt at this)." then "Finally, you could reboot the system, but you only have one shot at this
@@ -321,11 +179,6 @@ called the program break -- upward"); line 142 says "The end of the data segment
\keyword{program break}". Both are defensible historically, but stating both without comment will
confuse students.
### processes/processes.tex:928 — "environment variables cannot be read by an outside process"
On Linux an outside process with the right privileges can read `/proc/<pid>/environ`, so environment
variables are not a security boundary in the way this sentence implies. The intended point is probably
that they do not appear in `ps` output the way `argv` does. Worth an author correction.
### processes/processes.tex:834 — confusing claim about stdin/stdout/stderr after exec
"The operating system may open up 0, 1, 2 -- stdin, stdout, stderr, if they are closed after exec; most
of the time they leave them closed." Subject shifts from "operating system" to "they", and it is unclear
@@ -347,10 +200,6 @@ rewriting most of the paragraph, so I left it.
more to" invites a "than ..." that never arrives, and the claim itself (political reasons) is asserted
with no context a student could use.
### processes/processes.tex:812 — capitalization mismatch with the example
Prose says the example writes "Captain's Log"; the code at line 802 writes `"Captain's log"`. Trivial,
but the prose is quoting the program output.
### processes/processes.tex:185,340,881 — figures have captions but no alt text
`\includegraphics` of `address_space.eps`, `sleepsort_timing.eps`, and `fork_exec_wait.eps` carry only
`\caption{}`. For an accessible PDF these need real alternative descriptions (the sleepsort timing
@@ -366,44 +215,18 @@ diagram in particular carries information not present in its caption).
### malloc/malloc.tex:83 — "these limitations" has no antecedent
"An advanced discussion of these limitations is \href{...}{in this article}." The preceding sentence describes what `calloc` does; no limitations have been mentioned yet. A student cannot tell what limitations are meant. Also the linked host (locklessinc.com) may be dead — worth checking.
### ~~malloc/malloc.tex:85 — "calloc(x,y) is identical to calloc(y,x)"~~ WITHDRAWN — false positive
This item was raised and then **withdrawn on review**. The original claim
argued the book's statement was unsafe "once overflow checking is
considered", while simultaneously conceding that `n * size` overflow is
symmetric — which is self-contradictory. Multiplication is commutative, so
`calloc(x,y)` and `calloc(y,x)` request the same size and overflow at
exactly the same point. **The book is correct as written.** No action needed.
### malloc/malloc.tex:211 vs figure caption — "perfect-fit" vs "Best fit"
Prose says "A perfect-fit strategy finds the smallest hole"; the figure caption immediately below says "Best fit finds an exact match", and the rest of the chapter (and the Topics list) uses "Best Fit". Terminology inconsistency that could confuse a student; renaming is an editorial call.
### malloc/malloc.tex:237 — "don't need to replace the block"
"those placement strategies don't need to replace the block". Given the surrounding discussion of splitting and the following sentence about returning "the original block unbroken", this almost certainly should be "don't need to *split* the block". Changing it alters a technical claim, so flagging rather than fixing.
### malloc/malloc.tex:242 — "continuous block"
"it may be divided up in a way so a continuous block of that size is unavailable." Should almost certainly be "contiguous" — the standard term used elsewhere in this chapter (lines 8, 330). Flagged rather than fixed because it is a technical term.
### malloc/malloc.tex:266 — survey date vs citation
"a more rigorous survey conducted in 2005 \cite{10.1007/3-540-60368-9_19}". That DOI prefix (3-540-60368-9) is the 1995 Springer LNCS volume — Wilson, Johnstone, Neely & Boles, "Dynamic Storage Allocation: A Survey and Critical Review" (1995). The stated year appears wrong. Verify against malloc/malloc.bib.
### malloc/malloc.tex:290 — Fibonacci heaps claim
"Your heap could be represented with the max-heap data structure ... Using Fibonacci heaps, however, could be extremely inefficient." Fibonacci heaps have excellent amortized bounds; the claim as written is surprising and unexplained (presumably about constant factors / pointer overhead / cache behavior). Either justify or drop.
### malloc/malloc.tex:293 — next-fit definition is circular
"one is next-fit which is first fit on the next fit block" defines next-fit in terms of "the next fit block", which is undefined. A student reading only this sentence learns nothing. Needs a real one-line definition (resume the search where the last one stopped).
### malloc/malloc.tex:430-432 — broken quotation
The `quote` block ends: "...a multiple of 16 on 64-bit systems." For example, if you need to calculate how many 16 byte units are required, don't forget to round up." There is a stray closing double-quote mid-block, and the "For example..." sentence is the book's own commentary sitting inside the glibc quotation. Also the quoted text is self-contradictory ("always a multiple of eight on most systems"). Fixing requires deciding where the quotation actually ends, and possibly re-checking the glibc manual wording.
### malloc/malloc.tex:456-458 — free() sets is_free = 0
Prose says "A naive implementation would simply mark the block as unused. If we are storing the block allocation status in a bitfield, then we need to clear the bit", and the listing does `p->info.is_free = 0;`. With a field named `is_free`, marking a block unused means setting it to 1, not clearing it. Either the field name or the code is wrong. Not touched (code listing), but it reads as a genuine error rather than a deliberate teaching bug.
### malloc/malloc.tex:487 — incomplete sentence
"No more than 3 blocks will need to coalesce into a single block, and using a most recently used block scheme only one linked list entry." The second clause has no verb (presumably "...only one linked list entry needs to be updated"). Repairing it requires knowing the intended claim, so flagged rather than guessed.
### malloc/malloc.tex:646 — unit capitalization in exercise
"a new slab of 64kb ... allocating 1.5kb". Elsewhere the chapter uses KiB/KB consistently; "kb" reads as kilobits. Left alone since these are exercise numbers, but worth normalizing.
### Figures — no alt text
All figures (lines ~199-235, 306-319, 410-414, 475-479, 512-516, 530-534) use `\includegraphics` with a `\caption` only. The captions ("Malloc addition", "Free list good and bad coalesce") do not describe what the diagram shows, so a student using a screen reader or reading the text alone gets nothing. Accessibility improvement needs an author who knows the drawings.
@@ -411,70 +234,6 @@ All figures (lines ~199-235, 306-319, 410-414, 475-479, 512-516, 530-534) use `\
## threads
### threads/threads.tex:510 — code listing references undeclared `stack` (compile error)
The listing declares:
```
char *child_stack = malloc(STACK_SIZE);
```
but two lines later uses:
```
char *stack_top = stack + STACK_SIZE;
```
`stack` is never declared; the printed example does not compile. Open PR #212 proposes exactly this fix (`stack` -> `child_stack`). Left untouched because it is inside a `lstlisting`. Needs a human to land the fix / reconcile with the PR.
### threads/threads.tex:495-496 — comment says 8 KiB but the macro is 8 MiB
```
// 8 KiB stacks
#define STACK_SIZE (8 * 1024 * 1024)
```
`8 * 1024 * 1024` is 8 MiB. Inside a listing, so not edited; a human should decide whether the comment or the constant is wrong.
### threads/threads.tex:500 — missing semicolon in listing
```
puts("Hello Clone!")
```
No terminating semicolon; the example as printed does not compile.
### threads/threads.tex:501 — subject/verb in listing comment
```
// This share the same heap and address space!
```
"This share" should be "This shares" (or "These share"), but it is inside a listing, so not edited.
### threads/threads.tex:351,354,357 — bugs in the "thread-safe" solution listing
```
written = snprintf(buf, nbtytes, "%d : blah blah" , num);
```
`nbtytes` is a typo for `nbytes` (won't compile).
```
buf[nbytes] = '\0';
```
Writes one past the end of a buffer of `nbytes` bytes — a buffer overflow in a listing presented as "one valid solution". Also, `strncpy` already NUL-terminates here since "Unknown" is short.
`return written <= nbytes;` returns a truth value from a function whose name/usage suggests a byte count; worth a human check.
### threads/threads.tex:324 — "Race conditions aren't in our code."
```
Race conditions aren't in our code.
They can be in provided code.
```
As written the first sentence flatly denies race conditions exist in our code, which contradicts the preceding examples. The intended meaning is almost certainly "aren't only in our code". Meaning-changing, so left for a human.
### threads/threads.tex:205 — sentence fragment / duplicated "means"
```
@@ -484,22 +243,6 @@ Meaning that the same program can run multiple times and depending on how the ke
The second sentence is a fragment and repeats "means"; it also needs commas around the "depending on..." clause. Rewriting it is more than a mechanical fix.
### threads/threads.tex:32 — "processes" where "threads" is meant
```
\item When you want communication between the processes simplified
```
This bullet is in the list of reasons to prefer *threads*; "the processes" should probably be "threads". Technical wording, so not changed.
### threads/threads.tex:44 — missing qualifier changes the claim
```
It's easy to `free' the memory used by automatic variables because the program needs to change the stack pointer.
```
Intended sense is "because the program only needs to change the stack pointer". As printed the reasoning does not follow.
### threads/threads.tex:230-231 — confusing register description
```
@@ -517,14 +260,6 @@ We will instead treat i as a pointer and cast it by value.
The code passes the *value* of `i` cast to `void *`; "treat i as a pointer" is backwards, and "cast it by value" is not standard terminology. (The listing itself also uses `int data = ((int) ptr);`, which is implementation-defined on LP64 and normally warns.)
### threads/threads.tex:437 — complexity claim
```
The parallel algorithm runs in $O(\log^3(n))$ running time because the analysis assumes that we have a lot of cores.
```
The usual PRAM bound for parallel merge sort is $O(\log^2 n)$ (or $O(\log n)$ with the best merge). The exponent should be checked by a human.
### threads/threads.tex:3 — epigraph
```
@@ -545,30 +280,10 @@ What are a few things that threads share in a process? What are a few things tha
## synchronization
### Output string in prose does not match the code
`synchronization/synchronization.tex:44` — prose says a typical output is `\keyword{ARGGGH sum is <some number less than expected>}`, but the listing above (and the later corrected listing) both `printf("ARRRRG sum is %d\n", sum)`. One of the two spellings should win; I left the code alone per instructions.
### Loop counts in prose contradict the code
`synchronization/synchronization.tex:212-214` — the code loops `10000000` times, but the prose says "we lock and unlock the mutex a million times" and "add one million using an automatic (local) variable". Also line 213, "we could have added up twice!", is hard to parse — presumably "we could have just added the total twice" or similar. Needs an author decision on the intended wording/number.
### Broken initializer inside a listing (deliberate or not?)
`synchronization/synchronization.tex:243` — `m2 = = PTHREAD_MUTEX_INITIALIZER;` has a doubled `=` and would not compile. The listing's point is that two different mutexes protect the same variable, so the doubled `=` looks like an accidental typo rather than a deliberate bug, but it is inside a listing so I did not touch it.
### Mutex description may be misleading
`synchronization/synchronization.tex:235-236` — "If a mutex is locked, the other threads will continue. It's only when a thread attempts to lock a mutex that is already locked, will the thread have to wait." The second sentence is ungrammatical (a mixed "It is only when… that…" / "Only when… will…" construction). Rewording touches a technical claim, so I left it.
### Wrong identifier in prose
`synchronization/synchronization.tex:311` — "both threads would read \keyword{m\_locked} as zero" refers to the field written `m->locked` in the listing above. Identifier, so not changed.
### Factually wrong sentence about mutex initialization
`synchronization/synchronization.tex:354` — "We set the state of the mutex to unlocked and set the owner to locked." The code sets `mtx->owner = UNASSIGNED_OWNER`. "set the owner to locked" is meaningless; likely should be "unassigned".
### Description of weak vs strong CAS is garbled and arguably backwards
`synchronization/synchronization.tex:393` — "there are two versions to these atomic functions a \emph{strong} and a \emph{weak} part, strong guarantees the success or failure while weak may fail even when the operation succeeds." Missing punctuation, "part" is the wrong noun, and "may fail even when the operation succeeds" is a confusing way to state spurious failure (weak may fail even when the comparison succeeded). Needs an author rewrite.
@@ -593,34 +308,10 @@ What are a few things that threads share in a process? What are a few things tha
`synchronization/synchronization.tex:855-877` — the listing is labelled `// Sketch #1` and is syntactically broken (a `push` nested inside `pop`, unbalanced braces), and the very next paragraph starts "Sketch \#2 has implemented the \keyword{post} too early." Sketch #1 is never discussed. Reads like a missing paragraph.
### Typo'd identifier in a "correct" listing
`synchronization/synchronization.tex:935` — the fixed implementation calls `sem_init(&sremains, 0, SPACES)` while everything else uses `sremain`. As presented as the correct solution, this would not compile. In-listing, so not changed.
### Forward/backward cross-reference confusion
`synchronization/synchronization.tex:784-785` — "how would we fix the problems with condition variables? Try it out before you look at the code in the previous section." The condition-variable code being referenced follows immediately below, not in a previous section. Similarly line 810, "Does the following solution work?", appears *after* the listing it refers to.
### Wrong type name in the semaphore struct
`synchronization/synchronization.tex:1238` — `pthread_condition_t cv;` should be `pthread_cond_t`. Inside a listing, so not changed, but it will not compile as shown.
### "Three actions" list only has two ordinals
`synchronization/synchronization.tex:1610-1612` — "performs \emph{Three} actions. Firstly, it atomically unlocks the mutex and then sleeps … Thirdly, the awoken thread must re-acquire the mutex lock". "Secondly" is missing; either split the first sentence or renumber.
### Peterson's solution walkthrough is a fragment and possibly wrong
`synchronization/synchronization.tex:1204-1206` — "If thread \#2 has set turn to 2 and is currently inside the critical section. Thread \#1 arrives, \emph{sets the turn back to 1} and now waits until thread 2 lowers the flag." The first sentence is a fragment, and in the pseudocode above each thread sets `turn = other_thread_id`, so thread #2 would set turn to 1 and thread #1 would set it to 2. Technical — not touched.
### Bounded-wait definition is awkward
`synchronization/synchronization.tex:1034` — "A thread/process cannot get superseded by another thread infinite amounts of time." Probably "an infinite number of times". Left because the fix is a judgement call on the intended definition.
### Shared-mutex example uses undeclared variable names
`synchronization/synchronization.tex:2097-2125` — the listing declares `pthread_mutex_t *mutex` and `pthread_mutexattr_t attr`, then `main` uses `pmutex` and `attrmutex` throughout, and `write_string` locks `mutex` (never assigned). The code as printed cannot compile and the two halves do not share a mutex. It also ignores `write` returning -1 in the retry loop. Needs an author fix.
### Garbled question
`synchronization/synchronization.tex:2267` — "How might the above be a producer consumer problem be used in the above section?" Doubled "be" and duplicated "the above"; the intended question is unclear.
@@ -637,9 +328,6 @@ What are a few things that threads share in a process? What are a few things tha
## deadlock
### deadlock/deadlock.tex:36-40 and 42-66 — cycle detection presented as sufficient for deadlock, and `isCyclic` conflates "visited" with "on the current DFS path"
The text says "If there is a cycle in the Resource Allocation Graph and each resource in the cycle provides only one instance, then the processes will deadlock", then presents pseudocode whose only check is `if (this graph has been visited)`. Marking a node visited permanently and returning true on any revisit reports a cycle for a re-encountered node that is merely already finished (a DAG with a diamond), rather than one on the current recursion stack. A correct DFS cycle check needs a separate on-stack/in-progress marker. **This overlaps with open PR #195**, which proposes replacing this algorithm; per instructions I did not touch it. A human should decide between fixing the marker logic here and adopting #195.
### deadlock/deadlock.tex:79 — "necessary *and* sufficient conditions ... non-zero probability"
"There are four \emph{necessary} and \emph{sufficient} conditions for deadlock -- meaning if these conditions hold then there is a non-zero probability that the system will deadlock at any given iteration." Sufficiency normally means deadlock *does* occur, not that it *may* occur with non-zero probability; the gloss contradicts the term. The Coffman conditions are standardly stated as necessary (and sufficient only for single-instance resources). Needs an author decision, not a copy-edit.
@@ -652,9 +340,6 @@ Not a sentence, and it introduces the bulleted transition rules. Probably intend
### deadlock/deadlock.tex:123-129 — proof of the reverse direction
Line 124 begins "More formally, this system is deadlocked means if $\exists t_0, ...$" (missing a word/"that"), and line 129 claims "Hold and wait simply proves the condition that from this point onward, the system will not change, which is all the conditions that we needed to show." The logic of what is assumed versus proved is hard to follow, and line 107 says "let us build a system with the three requirements not including circular wait" while the argument then uses circular wait (lines 127-128). This is a substantive correctness question about the proof, not a wording issue.
### deadlock/deadlock.tex:137 — pen/paper contradiction example is wrong
"a student would have to be waiting on a pen while holding the paper and the other waiting on a pen and holding the paper." Both students are described identically. For the contradiction to work the second student must be waiting on the *paper* while holding the *pen*. This is a factual error in the example; I did not silently correct it because the intended assignment of items should be confirmed against the ordering rule stated on line 135.
### deadlock/deadlock.tex:145 — "philosopher" appears before the Dining Philosophers section
The livelock paragraph continues the pen-and-paper students example but says "if the philosopher picks up the same device again and again". Dining Philosophers is not introduced until section 4 (line 181). A student reading linearly has no referent yet.
@@ -676,12 +361,6 @@ Ungrammatical inside a proof; the intended sense is probably "acting under the p
### deadlock/deadlock.tex:378-383 — Dijkstra proof reduction
Line 378 "If the last philosopher $p_{n-1}$ holds the first lock meaning the previous philosopher $p_{n-2}$ is waiting on $r_{n-1}$ meaning $r_{n-2}$ is available" is a run-on with no main verb, and line 380 concludes "we now have $n$ resources but only $n-1$ philosophers" without stating which philosopher was removed. Also line 379 uses "her" while the rest of the passage uses "he/his". The substance of the reduction needs an author's check.
### deadlock/deadlock.tex:390 — figure caption does not match its section
The figure at the end of the Dijkstra / Partial Ordering subsection is captioned "Stalling solution partial deadlock", referencing Stallings' solution from the previous subsection. The image file is `dining_partial.eps`, so the caption looks like a copy-paste error from line 348 ("Stalling solution almost deadlock"). Also note the chapter spells the author "Stallings'" in the section title but "Stalling" in both captions.
### deadlock/deadlock.tex:37 — bare `\ref` without a prefix
"then processes 1 and 2 will be deadlocked \ref{ragfigure}." renders as "...deadlocked 1.1." with no "Figure" word and no punctuation cue. Should probably be "(see Figure~\ref{ragfigure})". Left alone since it touches a cross-reference.
### deadlock/deadlock.tex:24-29, 68-72, 191-195, 227-231, 268-272, 303-307, 345-349, 387-391 — figures have no alt text
Eight `\includegraphics` calls, none with alt text; several carry load-bearing content (the deadlock cycle, the livelock time evolution, the arbitrator diagram). Only three of the eight are referenced from the prose at all (`ragfigure` is the sole `\label`), so a reader relying on a screen reader loses the content entirely. Needs an accessibility decision at the book level.
@@ -689,22 +368,11 @@ Eight `\includegraphics` calls, none with alt text; several carry load-bearing c
## ipc
### ipc/ipc.tex:130 — Sentence truncated mid-thought, number missing
> "Meaning we need roughly
> With $2^{52}$ entries, that's $2^{55}$ bytes (roughly 40 petabytes)."
"Meaning we need roughly" is a dangling fragment with the value missing. Also the arithmetic needs a human check: 52-bit entries rounded to 8 bytes x $2^{52}$ entries = $2^{55}$ bytes = 32 PiB (~36 PB), not "roughly 40 petabytes". Fixing requires deciding the intended number, so I left it.
### ipc/ipc.tex:184 — Multi-level page table size arithmetic is confusing/possibly wrong
> "shrunk from 4MiB for the single-level implementation to three page tables of memory or 2KiB for the top-level and 4KiB for the two intermediate levels of size 10KiB."
Ungrammatical and technically muddled: it says "two intermediate levels" when the surrounding text (line 186-188) describes two *sub-tables*, and "4KiB for the two intermediate levels" reads as 4KiB total while 2+4+4=10KiB implies 4KiB each. A human should restate this sentence.
### ipc/ipc.tex:204 — "cache coherence" appears to be the wrong term
> "if a program has inadequate cache coherence, the address will be missing in the TLB"
The intended concept is locality of reference / poor TLB hit rate; "cache coherence" means something else entirely. Technical wording, so not silently changed.
### ipc/ipc.tex:217 — "read and write" in MMU pseudocode
> "get the physical frame from the TLB and perform the read and write."
@@ -720,22 +388,6 @@ Run-on with words apparently missing ("provide the address" is unattached) and n
"can access" has no object (access what — that page?). Needs an author to complete.
### ipc/ipc.tex:331 — Possibly outdated/overly narrow claim about mmap
> "Currently it only supports regular files and POSIX \keyword{shmem}"
Linux mmap also supports anonymous mappings (used later in this very chapter at line 415 with MAP_ANONYMOUS), device files, hugetlbfs, etc. Technical claim — flagged, not changed.
### ipc/ipc.tex:335-341 — PROT_* flags attributed to the wrong mmap argument
> "The first option is that the \keyword{flags} argument of mmap can take many options."
The list that follows starts with PROT_READ / PROT_WRITE / PROT_EXEC / PROT_NONE, which belong to the `prot` argument, not `flags`; only MAP_SHARED / MAP_PRIVATE are `flags`. Also line 339 says "If this is supplied and \keyword{PROT\_NONE} is also supplied, the latter wins" — PROT_NONE is 0, so OR-ing it changes nothing; that claim looks wrong. Both are technical statements about mmap.
### ipc/ipc.tex:375,385,389-398 — Inconsistent variable name in the mmap walkthrough
The text and code define `page_offset` (line 375) but the argument list (line 385) says "pa\_offset" and the `munmap` call (line 398) uses `pa_offset`, which would not compile as written. Inside listings, so not touched.
### ipc/ipc.tex:495-503 — Example command and its explanation disagree on the filename
The command is `tee dirents` but the prose says "\keyword{tee} outputs the contents to the file \keyword{dir\_contents}". One of the two needs to change; a human should pick which.
### ipc/ipc.tex:636 — Badly broken sentence about pipe writes
> "If a process tries to write with some reader's read goes through, or fails -- partially or completely -- if the pipe is full."
@@ -746,17 +398,6 @@ Words are clearly missing/garbled; the intended meaning (a write succeeds if the
Probably "a special byte". Left as-is since the chapter deliberately mixes second person elsewhere.
### ipc/ipc.tex:775-780 — fdopen example listing has real bugs
`open("mydata.txt", "w", O_CREAT, ...)` passes a string as the flags argument (should be `O_WRONLY | O_CREAT`), and the listing is missing its closing `return`/`}`. Inside a listing, so untouched, but this one does not look like a deliberate teaching bug.
### ipc/ipc.tex:859 — Reader/writer roles appear reversed
> "echo calls \keyword{open(..,\ O\_RDONLY)} but that blocks until cat calls \keyword{open(..,\ O\_WRONLY)}"
In the example above it, `echo Hello > fifo` is the *writer* (O_WRONLY) and `cat fifo` is the *reader* (O_RDONLY). The two look swapped. Technical claim, flagged rather than fixed.
### ipc/ipc.tex:890-901 — "Fine Pipe Access Pattern" table attributes print() to the wrong process
Time 4 lists "print() \& exit()" under Process 1, but Process 1 already did "close() \& exit()" at Time 3 and it is Process 2 that prints. Looks like a column error.
### ipc/ipc.tex:955,980 — "this"/"That quirk" with no antecedent
Section "Determining File Length" opens "using fseek and ftell is a simple way to accomplish this" (no prior referent), and section "Use stat instead" opens "This only works on some architectures and compilers. That quirk is that longs only need to be 4 Bytes big" — the quirk is named only after it is referred to. Reads as if an introductory sentence was lost.
@@ -779,14 +420,6 @@ The line introducing the shared example process list is just `Unless otherwise s
There is no `\ref`/`\label` here, and "the section conceptually scheduling" does not read like an actual section title. A human should confirm the target exists and ideally replace this with a real `\ref{}`. The sentence also has no terminal period, but I left it since the whole line may be rewritten.
### "time quanta" used as a singular noun
`scheduling/scheduling.tex:284-285`
> The maximum amount of time that a process can execute before being returned to the ready queue is called the time quanta.
> As the time quanta approaches infinity, ...
"quanta" is plural; the singular is "quantum" (and the chapter itself writes `Quantum = 1000ms` at line 307). Changing it is arguably a terminology decision for the course, so I left both occurrences alone. (I did fix "approaches to infinity" -> "approaches infinity" on line 285.)
### Stray capitalization "Convoy Behind them"
`scheduling/scheduling.tex:120`
@@ -794,13 +427,6 @@ There is no `\ref`/`\label` here, and "the section conceptually scheduling" does
"Behind" is capitalized mid-sentence for no apparent reason. It may be deliberate emphasis in this book's informal voice, so I left it. Also note "Convoy effect" (line 260) vs "Convoy Effect" (line 57) vs "convoy effect" (lines 118, 120, 262, 354) are inconsistently capitalized throughout.
### PSJF worked example: preemption description may not match the figure
`scheduling/scheduling.tex:213-220`
> Then P1 comes in at 1000ms, P2 runs for 2000ms, so our scheduler preemptively stops P2, and lets P1 run all the way through.
Two things a human should check against `scheduling/drawings/psjf.eps`: (a) the text says PSJF compares *total* runtimes (line 190), and P1 (1000ms) is strictly shorter than P2 (2000ms), so "the times are equal" / "completely up to the algorithm" (line 216) looks wrong here — the tie case is P4 vs P5, not P1 vs P2; (b) line 218 says "since the runtimes are equal to P5, the scheduler stops P5 and runs P4", but P4 is 4000ms and P5 is 5000ms, which are not equal. Per instructions I did not alter the example's logic or numbers.
### Figures have captions but no alt text
`scheduling/scheduling.tex:146-150, 193-197, 237-241, 287-291`
@@ -817,14 +443,6 @@ The point being made is about a process of *higher* priority than the running on
## networking
### Outdated IPv4/IPv6 adoption statistics — networking/networking.tex:64, 73
"Even as of 2018, IPv4 still dominates Internet traffic, but Google reports that 24 countries now supply 15\% of their traffic through IPv6" and "However, little web traffic is IPv6 based on comparison as of 2018". These 2018 figures are badly stale (Google's IPv6 adoption is far higher now). Needs a human to refresh the numbers and the `\cite{internet_society_2018}` reference.
### "The world ran out of IP addresses a while ago" — networking/networking.tex:93
Loose claim; IANA exhausted the free pool in 2011 and RIRs at various later dates. A human may want to make this precise. Also line 96: "these addresses are leased not bought" — IPv4 addresses are also allocated/leased rather than owned, so the stated contrast with IPv4 is questionable.
### Garbled IPv4 address-splitting sentence — networking/networking.tex:70
"Conceptually the source and destination addresses can be split into two: a network number the upper bits and lower bits represent a particular host number on that network." The sentence has no working structure and a student cannot extract the network/host split from it. Rewriting requires deciding what was meant, so I left it.
@@ -849,10 +467,6 @@ Unclear as written — presumably means high-performance / error-tolerant code s
"RFC 7231 has the most current specifications on the most common HTTP method today". RFC 7231 was obsoleted by RFC 9110 (HTTP Semantics, 2022), and the chapter's examples are all HTTP/1.0 while HTTP/1.1 and HTTP/2/3 dominate. A human should decide how much to update. Also line 553, "the HTTP/1.0 method" should probably be "protocol"/"version".
### Level-triggered epoll described in terms of the wrong call — networking/networking.tex:1286
"Level triggered means that while the file descriptor has events on it, it will be returned by epoll when calling the ctl function." Events are returned by `epoll_wait`, not `epoll_ctl`. Same issue at line 1291: "ctl will reset the state to zero events".
### Vague claim about epoll vs select — networking/networking.tex:1214
"There are reasons to use epoll over select but due to interface, there are fundamental problems with doing so." Ungrammatical and the point is unrecoverable — is the problem with select's interface or epoll's? Needs the author.
@@ -873,18 +487,10 @@ You send *packets*, not sockets. Likely "to send data over UDP sockets". I could
"Writing stub code by hand is painful, tedious, error-prone, difficult to maintain and difficult to reverse engineer the wire protocol from the implemented code." The final clause does not attach to the list.
### Bug in the RPC marshaling listing — networking/networking.tex:1330-1336
`int getHighScore(char* game)` but the body does `asprintf(&buffer,"getHiscore(%s)!", name);` — `name` is undeclared (should be `game`), and the comment says `'getHiscore'` while the function is `getHighScore`. Also `read(fd, buffer, sizeof(buffer))` uses `sizeof` on a `char*`. In a code listing, so possibly deliberate — but if not, students will copy it.
### Comma splice left as-is — networking/networking.tex:1319
"To marshal a linked list, it is unnecessary to send the link pointers, stream the values." Reads as a splice; the fix ("instead, stream the values") is a wording choice so I left it.
### "host-ordered ordering" — networking/networking.tex:333
"convert network ordered byte values to host-ordered ordering". Redundant/garbled; the natural fix is "host byte ordering" but it borders on style, so I left it.
### Figures have no alt text — networking/networking.tex:88-92, 236-240
Both `\includegraphics` figures (`ipv6_datagram.eps`, `tcp_header.eps`) carry only captions ("IPv6 Datagram divisibility", "Extra: TCP Header Specification") and no textual description. The IPv6 caption in particular does not explain what the diagram shows. Accessibility issue for screen-reader users.
@@ -893,32 +499,6 @@ Both `\includegraphics` figures (`ipv6_datagram.eps`, `tcp_header.eps`) carry on
## filesystems
### Garbled enumerate item in the chmod 755 example — filesystems.tex:693
> "\item r + w + x = digit * user has 4+2+1, full permission"
The "= digit \*" fragment looks like a mangled sub-bullet or a lost line break; the item also mixes the general rule and the user-specific case. The following two items ("group has...", "all users have...") are fine. Needs an author to decide the intended wording.
### `readdir` thread-safety advice appears inverted — filesystems.tex:370-371
> "\keyword{readdir} is not thread-safe! You shouldn't use the re-entrant version of the function."
The natural advice after "not thread-safe" is to *use* the re-entrant version (`readdir_r`), or to explain why `readdir_r` is deprecated in glibc and locking is preferred instead. As written the two sentences fight each other. Needs an author decision (the modern answer is "don't use `readdir_r`, use per-directory-stream locking", which the next sentence hints at but never says).
### Symlink `cat` example is internally inconsistent — filesystems.tex:425-446 (verbatim block)
`cat file1.txt` prints `file1!` but `cat file2.txt` prints `I'm file1!` even though `file2.txt` is a symlink to `file1.txt` and should print identical contents. Also `$ readlink myfile.txt` returns `file2.txt`, but `myfile.txt` is never created anywhere in the example — presumably it should be `readlink file2.txt` returning `file1.txt`. Inside a verbatim block, so not touched.
### ISO filename changes between mount example steps — filesystems.tex:950 vs 955 and 965
The download/mount commands use `archlinux-2015.04.01-dual.iso`, but the following prose and the `mount | grep arch` output use `archlinux-2014.11.01-dual.iso`. Filenames, so not touched — an author should pick one.
### `/dev/null/` with a trailing slash — filesystems.tex:811
> "Bytes sent to \keyword{/dev/null/} are never stored"
Every other mention is `/dev/null`. The trailing slash would actually fail (`/dev/null` is not a directory). Identifier, so not touched.
### "three times as slow" claim for indirection — filesystems.tex:226
> "This is three times as slow for reading between blocks, due to increased levels of indirection."
@@ -955,20 +535,10 @@ This should almost certainly be the 2nd *data block* (the inode is a single obje
Not a grammatical sentence, and it is unclear how it differs from the next question ("offset is greater than the length of the original file"). Left alone because the intended meaning is genuinely ambiguous.
### "0 being the inode root" — filesystems.tex:1164
> "The nth bit is set if the nth inode -- $0$ being the inode root -- is being used."
Almost certainly means "the root inode". Left alone in case "inode root" is intentional terminology for the minixfs model used here.
### Figure has no alt text — filesystems.tex:1174-1178
`\includegraphics{filesystems/images/sample_file.png}` with caption "Sample file filling up". The whole "Simple Filesystem Model" section (file size bounds, reads, writes) is written entirely against this image — a student using a screen reader, or reading the text alone, cannot follow any of the worked calculations. Adding a textual description of the inode's block pointers would fix this.
### `\begin {itemize}` with a stray space — filesystems.tex:1260
Compiles fine in LaTeX, but is inconsistent with every other environment in the file. Left alone as it is not a language issue.
### Fragment in "Writing to directories" — filesystems.tex:1268-1270
> "If we pretend that the example above is a directory. We know that we will be adding at most one directory entry at a time. Meaning that we have to have enough space for one directory entry in our data blocks."
@@ -979,36 +549,12 @@ Two sentence fragments in a row ("If we pretend..." with no main clause, and "Me
## signals
### signals.tex:219 — `\keyword(...)` uses parentheses instead of braces
`\item We execute \keyword(func("Hello"))` — the macro argument delimiters are wrong (compare line 222, `\keyword{func("World")}`). Needs a human to confirm intended markup rather than a blind edit. Note also that lines 220 and 223 use bare `strcmp(...)` in prose without `\keyword{}`, inconsistent with the surrounding text.
### signals.tex:256 — `\keyword{\!pleaseStop}` likely wrong escape
The sentence reads "The expression `\keyword{\!pleaseStop}` doesn't get changed in the body of the loop". `\!` is a negative thin space (math mode); in text mode it is not the `!` operator the sentence means. Probably should be a literal `!`, but the correct escaping inside `\keyword` depends on how that macro is defined, so a human should decide.
### signals.tex:468 — "any signal thread"
"A signal then can be delivered to any signal thread that is willing to accept that signal." "signal thread" looks like a typo for "single thread" or just "thread", but the two readings say different things about delivery semantics, so I did not guess.
### signals.tex:7 — "Sometimes, a program can choose to ignore events which is supported."
Circular/confusing as written; it is unclear whether the point is that ignoring is a supported disposition, or that only some signals may be ignored (SIGKILL/SIGSTOP cannot). A student would benefit from the caveat being stated explicitly here.
### signals.tex:16 — chapter promise not delivered
"This chapter will go over how to read information from a process that has either exited or been signaled." Nothing in this chapter covers `wait`/`waitpid` status macros (`WIFSIGNALED`, `WTERMSIG`, ...); that material lives in the processes chapter. Either the sentence or a cross-reference needs updating.
### signals.tex:58-62 — figure has no alt text
`\includegraphics{signals/drawings/signal_lifecycle.eps}` with caption "Signal lifecycle diagram" only. The caption does not convey the lifecycle content to a reader using a screen reader, and the surrounding text (line 56, "As a flowchart") does not describe it either.
### signals.tex:312, 355 — sample code missing statement terminators
`sigaction(SIGALRM, &sa, NULL)` and `sigprocmask(SIG_SETMASK, &set, &oldset)` both lack a semicolon; line 214 `printf("%s\n", buffer)` likewise, and `func` at 210 has no `return`. Left untouched per the code-listing rule, but if these are not deliberate they will confuse students copying them.
### signals.tex:194-195 — pseudo-C in listing
`raise(int sig);` and `kill(getpid(), int sig);` use declaration syntax as if calls. Intentional shorthand or a mistake — a human should decide whether to write `raise(sig)` / `kill(getpid(), sig)`.
### signals.tex:188 — `killall -l firefox`
The comment says "kill a process by executable name", but `-l` on `killall` lists signal names (GNU) rather than killing by name; the plain `killall firefox` seems intended. Technical, so reported rather than fixed.
### signals.tex:198-200 — "You can't SIGKILL any process!"
Ambiguous: reads as "no process can be SIGKILLed" rather than the intended "you can't SIGKILL just any process". Also `man -s2 kill` is the Solaris/BSD form; on Linux it is `man 2 kill`. Both are judgment calls.
### signals.tex:262 — sig_atomic_t range claim
"can be as small as a \keyword{char} and only able to represent (-127 to 127) values" — a technical claim about limits (C requires at least SIG_ATOMIC_MIN/MAX coverage) that I did not want to touch.
@@ -1031,17 +577,6 @@ Possessive of a singular noun ending in s-sound written as `process'`; elsewhere
"In lieu" requires "of X"; the intended phrase is probably "In lieu of that" or "Instead". Left alone as it may be deliberate shorthand.
### Spectre code example does not match its explanation — security/security.tex:180-202
Several mismatches a student would trip on:
- `char *a[10];` then `for (int i = 10; i != 1; --i) { a[i] = calloc(1,1); }` writes `a[10]`, which is out of bounds, and never allocates `a[1]`.
- The text says "The first loop allocates 9 elements" and "The last element is `0xCAFE`", but the code sets `a[0] = 0xCAFE` — the *first* element, not the last.
- `a[0] = 0xCAFE;` assigns an integer to a `char *` (would not compile cleanly).
- The second loop `for (int i = 10; i != 0; --i, --j)` shadows the outer `i` declared on line 188 ("This will be in main memory"), so the comment about register vs memory placement does not apply to the loop variable actually used.
- Text says "For the first 9 iterations, the branch is taken", but the loop runs 10 iterations.
This is presented as illustrative pseudo-code, but the index arithmetic contradicts the prose in a way that will confuse readers. Needs an author rewrite, not a copy-edit.
### "each user has a certain set of permissions that they can do" — security/security.tex:227
Grammatically mismatched ("permissions ... do") and conflates capabilities with permissions. Suggest "a certain set of capabilities" / "set of actions they are permitted to perform", but the wording sits inside a technical definition, so leaving to a human.
@@ -1052,24 +587,10 @@ Grammatically mismatched ("permissions ... do") and conflates capabilities with
Unclear what "the incorrect answer" refers to — the DNS response, or the decision to trust it. Reads as a garbled sentence; needs the author's intent.
### Dated DHS/DNSSEC paragraph — security/security.tex:344-345
> "As of 2019, the United States Department of Homeland Security released a directive to switch all services from DNS to DNSSec ..."
Three issues for a human: (a) "As of 2019" is now seven years stale and the directive (ED 19-01) was actually about mitigating DNS infrastructure tampering, not a blanket "switch all services from DNS to DNSSec"; (b) DNSSEC is not a replacement for DNS, it is an extension that adds origin authentication — "switch from DNS to DNSSec" is misleading; (c) the following sentence, "This directive is an inherent flaw of the DNS system.", is a broken sentence — presumably "This directive is *a response to* an inherent flaw". I did not repair it because the correct rewrite depends on the intended claim. The URL itself was left untouched per instructions.
Also worth noting for a refresh: this section predates the now-widespread deployment of DNS-over-HTTPS/DNS-over-TLS, which line 348 ("DNS requests are sent as unsecured UDP packets") states without qualification.
### Review question 381 vs body text — security/security.tex:336 and 381
Line 336 already states "Distributed Denial of Service is the hardest form of attack to stop", which answers review question 10 ("Which is harder to defend against: Syn-Flooding or Distributed Denial of Service?") outright. Intentional? Possibly fine, but flagging.
### "HTTPs" capitalization — security/security.tex:320
> "a higher level protocol such as HTTPs"
Should be "HTTPS". Left alone since it is a protocol identifier and outside the low-risk list.
---
## review
@@ -1077,12 +598,6 @@ Should be "HTTPS". Left alone since it is a protocol identifier and outside the
### review/review.tex:135 — truncated bonus question
The item ends: "Bonus: How would you make this code more robust or able to cope with?" The sentence is cut off ("cope with" what — a long `mesg`? a `malloc` failure?). The same sentence also says "val as a double val", which looks like a duplicated word but might be intentional shorthand. Both need an author who knows the intended question; rewriting could change what is being asked.
### review/review.tex:150 — two questions merged by a Markdown-conversion artifact
Line reads: "Why should you check the return value of sscanf and scanf? \#\# Q 5.2 Why is `gets' dangerous?" The literal `\#\# Q 5.2` is a leftover Markdown heading, and it welds two separate questions into one `\item`. Splitting them into two items is an editorial/structural change, so leaving it to a human.
### review/review.tex:156-159 — question/sub-part structure is broken
"What mistake did the programmer make in the following code? Is it possible to fix it?" is followed by two bare lines "i) using heap memory?" and "ii) using global (static) memory?" which LaTeX will run together into one paragraph rather than rendering as a sub-list. Needs a real `enumerate` or line breaks — a formatting decision.
### review/review.tex:197-203 — question text is split across a code listing
"When would a trivial malloc implementation" / listing / "be acceptable?" I lowercased the stray capital "Be", but the sentence still reads oddly when the listing is set as a display block. A human may prefer to reword (e.g. "When would the trivial malloc implementation shown below be acceptable?").
@@ -1092,18 +607,6 @@ Item at 322 sets up the graph/`shortest`/`set_edge` scenario and ends at line 33
### review/review.tex:519 — chmod question sentence is ungrammatical
"...so that the owner can read, write, and execute permissions the group can read and everyone else has no access." The verb "can" does not fit "permissions", and there is no punctuation separating the owner clause from the group clause. I only fixed the missing spaces after the commas; the rest is a rewrite that touches what the question asks (intended answer is presumably `chmod 740`), so a human should word it.
### review/review.tex:566 — handshake named incorrectly
"What is the SYN ACK ACK-SYN handshake?" The TCP three-way handshake is SYN, SYN-ACK, ACK. As written the term is wrong/garbled. Fixing it changes the technical content of the question, so flagging rather than editing.
### review/review.tex:607-613 — multiple-choice answer may be missing/ambiguous
"Assuming a network has a 20ms One Way Transit Time... how much time would it take to establish a TCP Connection?" Options are 20/40/100/60 ms. The three-way handshake is 1.5 RTT = 60ms to complete, but a connection is often counted as usable after 40ms (client can send with the final ACK). Which option is intended needs the author.
### review/review.tex:517 — "double direction table" corrected to "double indirection table"
I treated this as a typo and fixed it, but flagging in case "direction" was deliberate shorthand. Also note the item asks "What is the minimum file size required to require a single indirection table?" — "required to require" is clumsy but I left it because rewording risks changing the question.
### review/review.tex:649 — signal that "can not be caught"
Original: "Give the name of a signal that can not be caught by a signal". I added "handler" to complete the sentence. If the intended answer is SIGKILL/SIGSTOP, the fuller phrasing "cannot be caught, blocked, or ignored" may be preferable — an author call.
### review/review.tex:513,515,585 — space before question mark
Three items have "notes.txt} ?" / "listen accept ?" with a space before the "?". Left alone as it may be a deliberate consequence of the `\keyword{}` macro spacing, but a human may want them tightened.
@@ -1111,33 +614,9 @@ Three items have "notes.txt} ?" / "listen accept ?" with a space before the "?".
## honors
### honors/tcp.tex:3 — chapter section is an unfinished fragment
The entire file is one truncated sentence: "TCP or the transmission control protocol handles the retransmission, acknowledgement, and flow of packets. Naturally a pre-requisite to this section is" — it stops mid-clause with no terminal punctuation. An author needs to finish the sentence and write the section; I cannot invent the missing content.
### honors/honors.tex:8-9 — tcp.tex is never included
`honors.tex` only does `\input{honors/kernel.tex}` and `\input{honors/containers.tex}`. `honors/tcp.tex` exists but is orphaned, so the "TCP In Depth" section never renders. Either finish and include it, or delete the file. Human decision.
### honors/containers.tex:23-27 — three empty subsections
`\subsection{Linux Namespaces}`, `\subsection{Building a container from scratch}`, and `\subsection{Containers in the wild: Software distribution is a Snap}` have no body text at all. They will render as bare headings in the student-facing PDF.
### honors/containers.tex:2-3 — dated statistic
"around 20 billion devices connected to the internet in 2018" — an eight-year-old figure presented in the present tense ("As we enter an era..."). Needs refreshing or rephrasing as a historical reference.
### honors/kernel.tex:29 — technically questionable claim about traps
"System Calls use an instruction that can be run by a program operating in userspace that \textit{traps} to the kernel (by use of a signal) to complete the call." The parenthetical "(by use of a signal)" conflates a hardware trap/software interrupt (`syscall`/`int 0x80`) with a POSIX signal, which are unrelated mechanisms. This is exactly the kind of subtle technical claim I was told not to silently rewrite, but it reads as wrong and students in CS 341 have already learned what signals are.
### honors/kernel.tex:11-13 — Windows/Darwin sentence was structurally broken
Original text read as two fragments: "...the Windows kernel, which we won't talk about too much in this chapter." followed by a new line beginning "or \keyword{Darwin}, the UNIX-like kernel for macOS...". I joined them with a comma (minimal fix), but the result now reads as "we won't talk about Windows or Darwin", which may not be the intended meaning — the original may have lost a clause such as "Others may have used XNU or Darwin". Please confirm the intended sentence.
### honors/kernel.tex:17-19 — micro-kernel description may be inaccurate
"A micro-kernel ... provides the bare-minimum functionality that a kernel needs. This involves setting up important device drivers, the root filesystem, paging..." Putting device drivers and the root filesystem *inside* the micro-kernel contradicts the next sentence, which says drivers and filesystems live outside as separate servers. Needs an author to reconcile.
### honors/kernel.tex:69-73 — commented-out TODO outline left in source
Five commented lines ("How to add a new system call?", "What is a kernel module?", "Example kernel module.", etc.) mark unwritten material. Flagging so the chapter's incompleteness is visible.
### honors/kernel.tex:13 — inconsistent capitalization of a project name
"\keyword{zircon}" is lowercase while "\keyword{GNU HURD}" is uppercase; the Fuchsia kernel is normally written "Zircon". I did not change it because it is inside a `\keyword{}` identifier.
---
## appendix
@@ -1164,20 +643,6 @@ Ungrammatical count/mass mismatch; likely intended "a few filesystem hardware te
Section `\subsection{Implementing Software Mutex}` opens with a bare line reading `Yes` before "With a bit of searching, it is possible to find it in production...". This looks like the answer to a question that was deleted (probably "Is Peterson's algorithm ever used in practice?"). The paragraph also has an unexplained "it". A human should restore the missing question or delete the line.
### appendix/appendix.tex:762 — double negative in a code comment
> `/* Even though the other thread is woken up it cannot not return */`
"cannot not return" should almost certainly be "cannot return". It is inside an `lstlisting`, so I left it untouched per instructions.
### appendix/appendix.tex:945-953 — Sequentially Consistent example does not match its explanation
The listing stores `x` = 1 then 0, and sets `y = 10`; the prose says:
> "This is because either the store happens before the if statement in thread 2 and y == 1 or the store happens after and x does not equal 2."
`y == 1` and `x does not equal 2` do not correspond to any value in the example (the values are `y == 10`, `x` in {0,1}). Also "Will never quit" is presumably "will never abort". This is a memory-model claim, so I did not touch it.
### appendix/appendix.tex:943 — Sequential consistency definition is hard to parse
> "This model says that any change that happens, all changes before it will be synchronized between all threads."
@@ -1190,18 +655,6 @@ Grammatically broken and technically imprecise (sequential consistency is about
The language name appears three times and the sentence has no terminal punctuation. Probably intended: "We'll talk about a language similar to C in terms of simplicity and design: Go (or golang)." Left alone because it is a rewrite, not a typo fix.
### appendix/appendix.tex:1139 — "variadic" used to mean "variable"
> "when $C > 1$ we say that the running times of the process are variadic"
"Variadic" means "taking a variable number of arguments"; the intended word is almost certainly "variable" or "highly variable". Same misuse appears at line 1387 ("a variadic, pyramid summation") and line 1444 ("this is variadic in size"). Terminology change — author's call.
### appendix/appendix.tex:1166 — "Conway and Al"
> "discovered by Conway and Al \cite{conway1967theory}"
This should be "Conway et al." I left it because it sits directly next to a `\cite{}` and could conceivably be an intentional (if odd) rendering.
### appendix/appendix.tex:1126 vs 1317 — inconsistent symbol for maximum run time
Line 1126 says "Let the maximum amount of time that a process runs be equal to $S$", but line 1317 says "$T$ is the maximum amount of time a process can run for". Meanwhile $S$ is used throughout as the *service time* random variable. A student following the derivations would be confused; needs the author to pick one symbol.
@@ -1218,13 +671,6 @@ Mixes third and second person and is missing words. Rewrite needed.
"Either" has no second alternative. Likely a dropped clause.
### appendix/appendix.tex:1416-1417 — IPv4 header field sizes are wrong
> "The first octet is the version number, either 4 or 6"
> "The next octet is how long the header is."
In IPv4 the version is 4 bits and IHL is the *other* 4 bits of the same first octet — not two separate octets. The subsequent items ("The next octet is various bit flags", "The next octet and half is fragment number") are likewise off: flags are 3 bits and fragment offset is 13 bits, sharing two octets. TTL and protocol are one octet each, which is correct. This whole enumeration needs a technical pass against the actual header layout.
### appendix/appendix.tex:1461 — orphan fragment in the routing list
> "These protocols are meant to be fast and more trusting because all computers, switches, and routers are part of an ISP.
@@ -1244,10 +690,6 @@ The sentence is garbled; I could not tell which "assuming" to drop or how the in
Not a grammatical sentence; probably intended "kqueue is descriptor-agnostic in the truest sense." Rewrite needed.
### appendix/appendix.tex:1512 — product name capitalization
"MacOs" should be "macOS". Left alone because it borders on an identifier/product name and the book may spell it this way elsewhere.
### appendix/appendix.tex:171-198, 1409-1413 — figures have no alt text
`\includegraphics` for `struct_clean.eps`, `struct_slop.eps`, and `ip_datagram.eps` have captions ("Six box struct", "IP Datagram divisibility") but no descriptive alternative text. The IP datagram figure in particular carries information not otherwise in the text. Accessibility decision for the author.
@@ -1265,9 +707,6 @@ Not a grammatical sentence; probably intended "kqueue is descriptor-agnostic in
### Mars Pathfinder paragraph: run-on and tense-mixing
`post_mortems/post_mortems.tex:74,76` — Line 74 mixes past and present ("The finder uses a single bus...", "if an interrupt happened ... and a task is running and a task is to be scheduled"). Line 76 is a comma splice: "The pattern that caused everything to start failing was the data collection thread starts writing to the bus, the information bus thread is waiting on the data." Fixing this properly means restructuring sentences, which is beyond a low-risk copy-edit. Note the classic name for this bug — priority inversion — is never stated, which is the one term a student would want.
### "Required Sections: Intro to C/Appending" — likely wrong section name
`post_mortems/post_mortems.tex:179` — Every other section lists real chapter names ("Intro to C", "Malloc", "Processes"). "Appending" looks like a typo for "Appendix" (the Shell Shock section uses "Appendix/Shell"). I did not guess. These "Required sections" lines are also plain text, not `\ref{}` cross-references, so nothing verifies they point at chapters that exist.
### Sentence fragment in the Sony rootkit section
`post_mortems/post_mortems.tex:156` — "What websites visited, what clicks or keys typed etc." has no verb. It reads as a deliberate telegraphic aside in an informal chapter, so I left it, but a human may want "What websites are visited, what clicks or keys are typed, etc."
@@ -1277,8 +716,5 @@ Not a grammatical sentence; probably intended "kqueue is descriptor-agnostic in
### No `\label{}` anywhere in the chapter
`post_mortems/post_mortems.tex` (whole file) — The chapter and its ~16 sections define no labels, so nothing elsewhere in the book can cross-reference an individual post-mortem. Structural, and a human's call.
### Straight quotes in the opening line
`post_mortems/post_mortems.tex:5` — `a big "why are we learning all of this"` uses straight ASCII quotes; line 41 correctly uses LaTeX `` ``cat'' ``. This renders as two right-facing quotes in the PDF. Trivially fixable but it is typography, so flagging rather than changing.
---
+2 -8
View File
@@ -1,6 +1,6 @@
\section{Containerization}
As we enter an era of unprecedented scale with around 20 billion devices
connected to the internet in 2018, we need technologies that help us develop and
We live in an era of unprecedented scale, with billions of devices
connected to the internet, so we need technologies that help us develop and
maintain software capable of scaling upwards. Additionally, as software increases
in complexity, and designing secure software becomes harder, we find that we
have new constraints imposed on us as we develop applications. As if that wasn't
@@ -19,9 +19,3 @@ a host machine, while isolating itself from other containers or processes on the
host. You may have encountered containers while working with technologies such
as \keyword{Docker}, perhaps the most well-known implementation of containers
out there.
\subsection{Linux Namespaces}
\subsection{Building a container from scratch}
\subsection{Containers in the wild: Software distribution is a Snap}
+5 -11
View File
@@ -10,14 +10,14 @@ We will mostly be focusing on the Linux kernel in this chapter, so please assume
As it stands, most of you are probably familiar with the Linux kernel, at least in terms of interacting with it via system calls.
Some of you may also have explored the Windows kernel, which we won't talk about too much in this chapter,
or \keyword{Darwin}, the UNIX-like kernel for macOS (a derivative of BSD).
Those of you who might have done a bit more digging might have also encountered projects such as \keyword{GNU HURD} or \keyword{zircon}.
Those of you who might have done a bit more digging might have also encountered projects such as \keyword{GNU HURD} or \keyword{Zircon}.
Kernels can generally be classified into one of two categories, a monolithic kernel or a micro-kernel. A monolithic
kernel is essentially a kernel and all of its associated services as a single program. A micro-kernel on the other hand
is designed to have a \textit{main} component which provides the bare-minimum functionality that a kernel needs. This
involves setting up important device drivers, the root filesystem, paging or other functionality that is imperative for
other higher-level features to be implemented. The higher-level features (such as a networking stack, other filesystems,
and non-critical device drivers) are then implemented as separate programs that can interact with the kernel by some
typically means scheduling, memory management and paging, and IPC, the functionality that is imperative for
other higher-level features to be implemented. The higher-level features (such as a networking stack, filesystems,
and device drivers) are then implemented as separate programs that can interact with the kernel by some
form of IPC, typically RPC. As a result of this design, micro-kernels have traditionally been slower than monolithic
kernels due to the IPC overhead.
@@ -26,7 +26,7 @@ We will devote our discussion from here onwards to focusing on monolithic kernel
\subsection{System Calls Demystified}
System Calls use an instruction that can be run by a program operating in userspace that \textit{traps} to the kernel (by use of a signal) to complete the call.
System Calls use an instruction that can be run by a program operating in userspace that \textit{traps} to the kernel (for example, the \keyword{syscall} instruction on x86-64, not a POSIX signal) to complete the call.
This includes actions such as writing data to disk, interacting directly with hardware in general or operations related to gaining or relinquishing privileges (e.g. becoming the root user and gaining all capabilities).
In order to fulfill a user's request, the kernel will rely on \keyword{kernel calls}.
@@ -66,9 +66,3 @@ GFP_ATOMIC - Allocation will not sleep. May use emergency pools. For example, us
You'll note that some flags are marked as potentially causing sleeps.
This tells us whether we can use those flags in special scenarios, like interrupt contexts, where speed is of the essence, and operations that may block or wait for another process may never complete.
%* How to add a new system call?
%* What are processes and threads according to the kernel?
%* How is a scheduler actually implemented in the kernel?
%* What is a kernel module?
%* Example kernel module.
-3
View File
@@ -1,3 +0,0 @@
\section{TCP In Depth}
TCP or the transmission control protocol handles the retransmission, acknowledgement, and flow of packets. Naturally a pre-requisite to this section is
+1 -1
View File
@@ -6,7 +6,7 @@ The C memory model is probably unlike most that you've seen before. Instead of a
In low-level terms, a struct is a piece of contiguous memory, nothing more.
Just like an array, a struct has enough space to keep all of its members.
But unlike an array, it can store different types. Consider the contact struct declared above.
But unlike an array, it can store different types. Consider the contact struct declared below.
\begin{lstlisting}[language=C]
struct contact {
+7 -6
View File
@@ -208,7 +208,7 @@ To have a library function parse input in addition to reading it, use \keyword{s
All of those functions will return how many items were parsed.
It is a good idea to check if the number is equal to the amount expected.
Also naturally like \keyword{printf}, \keyword{scanf} functions require valid pointers.
Instead of pointing to valid memory, they need to also be writable.
In addition to pointing to valid memory, they need to also be writable.
It's a common source of error to pass in an incorrect pointer value.
For example,
@@ -220,7 +220,7 @@ char type;
int ok = 2 == sscanf(line, "%c %d", &type, &data); // pointer error
\end{lstlisting}
We wanted to write the character value into c and the integer value into the malloc'd memory.
We wanted to write the character value into \keyword{type} and the integer value into the malloc'd memory.
However, we passed the address of the data pointer, not what the pointer is pointing to!
So \keyword{sscanf} will change the pointer itself.
The pointer will now point to address 10 so this code will later fail when free(data) is called.
@@ -249,12 +249,13 @@ Any behavior missing from the documentation, such as the result of \keyword{strl
\begin{itemize}
\item \keyword{int strlen(const char *s)} returns the length of the string.
\item \keyword{size\_t strlen(const char *s)} returns the length of the string.
\item \keyword{int strcmp(const char *s1, const char *s2)} returns an integer determining the lexicographic order of the strings.
If s1 were to come before s2 in a dictionary, then a -1 is returned.
If s1 were to come before s2 in a dictionary, then a negative value is returned.
If the two strings are equal, then 0.
Else, 1.
Else, a positive value.
Only the sign is guaranteed, so compare the result with 0 rather than with -1 or 1.
\item \keyword{char *strcpy(char *dest, const char *src)} Copies the string at \keyword{src} to \keyword{dest}.
\textbf{This function assumes dest has enough space for src otherwise undefined behavior}
@@ -334,7 +335,7 @@ tricky
const char *input = "0"; // or "!##@" or ""
char* endptr;
int saved_errno = errno;
errno = 0
errno = 0;
long int parsed = strtol(input, &endptr, 10);
if(parsed == 0 && errno != 0){
// Definitely an error
+1 -1
View File
@@ -3,7 +3,7 @@
C was developed by Dennis Ritchie and Ken Thompson at Bell Labs back in 1973 \cite{Ritchie:1993:DCL:155360.155580}.
Back then, we had gems of programming languages like Fortran, ALGOL, and LISP.
The goal of C was two-fold.
Firstly, it was made to target the most popular computers at the time, such as the PDP-7.
Firstly, it was made to target the most popular computers at the time, such as the PDP-11.
Secondly, it tried to remove some of the lower-level constructs (managing registers, and programming assembly for jumps), and create a language that had the power to express programs procedurally (as opposed to mathematically like LISP) with readable code.
All this while still having the ability to interface with the operating system.
It sounded like a tough feat.
+13 -12
View File
@@ -236,7 +236,7 @@ if (connect(...)) {
exit(-2);
}
// (1)
// (4)
if (connect(...)) {
exit(-1);
} else if (bind(..)) {
@@ -246,22 +246,22 @@ if (connect(...)) {
}
\end{lstlisting}
\item \keyword{inline} is a compiler keyword that tells the compiler it's okay to omit the C function call procedure and "paste" the code in the callee.
Instead, the compiler is hinted at substituting the function body directly into the calling function.
\item \keyword{inline} is a compiler keyword that tells the compiler it's okay to omit the C function call procedure and "paste" the code in the caller.
That is, the compiler is hinted at substituting the function body directly into the calling function.
This is not always recommended explicitly as the compiler is usually smart enough to know when to \keyword{inline} a function for you.
\begin{lstlisting}[language=C]
inline int max(int a, int b) {
return a < b ? a : b;
static inline int max(int a, int b) {
return a > b ? a : b;
}
int main() {
printf("Max %d", max(a, b));
// printf("Max %d", a < b ? a : b);
// printf("Max %d", a > b ? a : b);
}
\end{lstlisting}
\item \keyword{restrict} is a keyword that tells the compiler that this particular memory region shouldn't overlap with all other memory regions.
\item \keyword{restrict} is a keyword that tells the compiler that this particular memory region shouldn't overlap with any other memory regions.
The use case for this is to tell users of the program that it is undefined behavior if the memory regions overlap.
Note that memcpy has undefined behavior when memory regions overlap.
If this might be the case in your program, consider using memmove.
@@ -273,9 +273,9 @@ void add_array(int *a, int * restrict c) {
*a += *c;
}
int *a = malloc(3*sizeof(*a));
*a = 1; *a = 2; *a = 3;
add_array(a + 1, a) // Well defined
add_array(a, a) // Undefined
a[0] = 1; a[1] = 2; a[2] = 3;
add_array(a + 1, a); // Well defined
add_array(a, a); // Undefined
\end{lstlisting}
\item \keyword{return} is a control flow operator that exits the current function.
@@ -283,11 +283,11 @@ add_array(a, a) // Undefined
Otherwise, another parameter follows as the return value.
\begin{lstlisting}[language=C]
void process() {
int process() {
if (connect(...)) {
return -1;
} else if (bind(...)) {
return -2
return -2;
}
return 0;
}
@@ -348,6 +348,7 @@ sizeof(str2) //8 because it is a pointer
\begin{enumerate}
\item When used with a global variable or function declaration it means that the scope of the variable or the function is only limited to the file.
\item When used with a function variable, that declares that the variable has static allocation -- meaning that the variable is allocated once at program startup not every time the program is run, and its lifetime is extended to that of the program.
\item When used inside the brackets of an array parameter (C99), as in \keyword{void f(int a[static 10])}, it promises that the caller passes a pointer to at least that many elements.
\end{enumerate}
\begin{lstlisting}[language=C]
+5 -2
View File
@@ -23,12 +23,15 @@ if (answer = 42) { printf("The answer is %d", answer);}
The quick way to fix that is to get in the habit of putting constants first.
This mistake is common enough in while loop conditions.
Most modern compilers disallow assigning variables a condition without parenthesis.
Most modern compilers warn about an assignment used as a condition unless it is wrapped in an extra set of parentheses (for example, \keyword{gcc -Wall} enables \keyword{-Wparentheses}).
\begin{lstlisting}[language=C]
if (42 = answer) { printf("The answer is %d", answer);}
\end{lstlisting}
With the constant first, the missing \keyword{=} becomes a compile error, because you cannot assign to \keyword{42}.
Once you notice the error, write \keyword{if (42 == answer)}.
There are cases where we want to do it.
A common example is getline.
@@ -49,7 +52,7 @@ time_t start = time();
\end{lstlisting}
The system function `time' actually takes as a parameter a pointer to some memory that can receive the time\_t structure or NULL.
The compiler fails to catch this error because the programmer omitted the valid function prototype by including \keyword{time.h}.
The compiler fails to catch this error because the programmer did not include \keyword{time.h}, so there is no valid function prototype to check the call against.
More confusingly this could compile, work for decades and then crash.
The reason for that is that time would be found at link time, not compile-time in the C standard library which almost surely is already in memory.
+2 -2
View File
@@ -112,8 +112,8 @@ printf("%s", ptr);
\end{lstlisting}
Notice how only 'EFGH' is printed.
Why is that? Well as mentioned above, when performing 'bna+=1' we are increasing the **integer** pointer by 1, (translates to 4 bytes on most systems) which is equivalent to 4 characters (each character is only 1 byte).
Because pointer arithmetic in C is always automatically scaled by the size of the type that is pointed to, POSIX standards forbid arithmetic on void pointers.
Why is that? Well as mentioned above, when performing 'bna+=1' we are increasing the \textbf{integer} pointer by 1, (translates to 4 bytes on most systems) which is equivalent to 4 characters (each character is only 1 byte).
Because pointer arithmetic in C is always automatically scaled by the size of the type that is pointed to, the ISO C standard forbids arithmetic on void pointers.
Having said that, compilers will often treat the underlying type as \keyword{char}.
Here is a machine translation.
The following two pointer arithmetic operations are equal.
-51
View File
@@ -1,51 +0,0 @@
\section{Topics}
\begin{itemize}
\tightlist
\item
C-strings representation
\item
C-strings as pointers
\item
char p{[}{]}vs char* p
\item
Simple C string functions (strcmp, strcat, strcpy)
\item
sizeof char
\item
sizeof x vs x*
\item
Heap memory lifetime
\item
Calls to heap allocation
\item
Dereferencing pointers
\item
Address-of operator
\item
Pointer arithmetic
\item
String duplication
\item
String truncation
\item
double-free error
\item
String literals
\item
Print formatting.
\item
memory out of bounds errors
\item
static memory
\item
file input / output. POSIX vs. C library
\item
C input output: fprintf and printf
\item
POSIX file IO (read, write, open)
\item
Buffering of stdout
\end{itemize}
+20 -18
View File
@@ -127,8 +127,7 @@ The remaining upper bits will be the page number (111100001111000011110000). Thi
We do have a problem with 64-bit operating systems.
For a 64-bit machine with 4KiB pages, each entry needs 52 bits.
Meaning we need roughly
With $2^{52}$ entries, that's $2^{55}$ bytes (roughly 40 petabytes).
Rounding each entry up to 8 bytes, with $2^{52}$ entries that's $2^{55}$ bytes (32 PiB, roughly 36 petabytes).
So our page table is too large.
In 64-bit architecture, memory addresses are sparse, so we need a mechanism to reduce the page table size, given that most of the entries will never be used.
We'll talk about this below. There is one last piece of terminology that needs to be covered.
@@ -201,7 +200,7 @@ To overcome this overhead, the MMU includes an associative cache of recently-use
This cache is called the TLB (``translation lookaside buffer'').
Every time a virtual address needs to be translated into a physical memory location, the TLB is queried in parallel to the page table.
For most memory accesses of most programs, there is a significant chance that the TLB has cached the results.
However, if a program has inadequate cache coherence, the address will be missing in the TLB, meaning the MMU must use the much slower page table translation.
However, if a program has poor locality of reference, the address will often be missing in the TLB, meaning the MMU must use the much slower page table translation.
\subsection{MMU Algorithm}
@@ -328,18 +327,18 @@ Page faults are so powerful because they let the operating system take control o
\keyword{mmap} does more than take a file and map it to memory.
It is the general interface for creating shared memory among processes.
Currently it only supports regular files and POSIX \keyword{shmem} \cite{mmap_2018}.
POSIX requires it to support regular files and POSIX \keyword{shmem} objects \cite{mmap_2018}; Linux also supports anonymous mappings that are not backed by any file, which we use later in this chapter.
Naturally, you can read all about it in the reference above, which references the current working group POSIX standard.
Some other options to note in the page will follow.
The first option is that the \keyword{flags} argument of mmap can take many options.
The \keyword{prot} and \keyword{flags} arguments of mmap can take many options: the \keyword{PROT\_*} values go in \keyword{prot} and the \keyword{MAP\_*} values go in \keyword{flags}.
\begin{enumerate}
\item \keyword{PROT\_READ} This means the process can read the memory. This isn't the only flag that gives the process read permission, however! The underlying file descriptor, in this case, must be opened with read privileges.
\item \keyword{PROT\_WRITE} This means the process can write to the memory. This has to be supplied for a process to write to a mapping. If this is supplied and \keyword{PROT\_NONE} is also supplied, the latter wins and no writes can be performed. The underlying file descriptor, in this case, must either be opened with write privileges or a private mapping must be supplied below.
\item \keyword{PROT\_EXEC} This means the process can execute this piece of memory. Although this is not stated in POSIX documents, this shouldn't be supplied with WRITE or NONE because that would make this invalid under the NX bit or not being able to execute (respectively).
\item \keyword{PROT\_NONE} This means the process can't do anything with the mapping. This could be useful if you implement guard pages in terms of security. If you surround critical data with many more pages that can't be accessed, that decreases the chance of various attacks.
\item \keyword{MAP\_SHARED} This mapping will be synchronized to the underlying file object. The file descriptor must've been opened with write permissions in this case.
\item \keyword{PROT\_WRITE} This means the process can write to the memory. This has to be supplied for a process to write to a mapping. The underlying file descriptor, in this case, must either be opened with write privileges or a private mapping must be supplied below.
\item \keyword{PROT\_EXEC} This means the process can execute this piece of memory. Although this is not stated in POSIX documents, this shouldn't be supplied with WRITE because that would make this invalid under the NX bit.
\item \keyword{PROT\_NONE} This means the process can't do anything with the mapping. Its value is zero, so it can't be combined with the other \keyword{PROT\_*} values to take permissions away. This could be useful if you implement guard pages in terms of security. If you surround critical data with many more pages that can't be accessed, that decreases the chance of various attacks.
\item \keyword{MAP\_SHARED} This mapping will be synchronized to the underlying file object. If the mapping is also writable (\keyword{PROT\_WRITE}), the file descriptor must've been opened with write permissions.
\item \keyword{MAP\_PRIVATE} This mapping will only be visible to the process itself. Useful to not thrash the operating system.
\end{enumerate}
@@ -382,7 +381,7 @@ Then, we make the call to mmap, here is the order of arguments.
\item PROT\_READ, we want to read the file
\item MAP\_PRIVATE, tell the OS, we don't want to share our mapping
\item fd, object descriptor that we refer to
\item pa\_offset, the page aligned offset to start from
\item page\_offset, the page aligned offset to start from
\end{enumerate}
\begin{lstlisting}[language=C]
@@ -395,7 +394,7 @@ After, we have to unmap the file and close the file descriptor to make sure othe
\begin{lstlisting}[language=C]
write(1, addr + offset - page_offset, length);
munmap(addr, length + offset - pa_offset);
munmap(addr, length + offset - page_offset);
close(fd);
\end{lstlisting}
@@ -500,7 +499,7 @@ First, it lists the current directory.
The -1 means that it outputs one entry per line.
The \keyword{cut} command then takes everything before the first period.
\keyword{sort} sorts all the input lines, \keyword{uniq} makes sure all the lines are unique.
Finally, \keyword{tee} outputs the contents to the file \keyword{dir\_contents} and the terminal for your perusal.
Finally, \keyword{tee} outputs the contents to the file \keyword{dirents} and the terminal for your perusal.
The important part is that bash creates \textbf{5 separate processes} and connects their standard outs/stdins with pipes. The trail looks something like this.
\begin{figure}[H]
@@ -768,15 +767,18 @@ If the program already has a file descriptor, it can `wrap' it into a FILE point
#include <sys/types.h>
#include <sys/stat.h>
#include <fcntl.h>
#include <stdio.h>
int main() {
char *name="Fred";
int score = 123;
int filedes = open("mydata.txt", "w", O_CREAT, S_IWUSR | S_IRUSR);
int filedes = open("mydata.txt", O_WRONLY | O_CREAT | O_TRUNC, S_IWUSR | S_IRUSR);
FILE *f = fdopen(filedes, "w");
fprintf(f, "Name:%s Score:%d\n", name, score);
fclose(f);
fclose(f); // also closes filedes
return 0;
}
\end{lstlisting}
For writing to files, this is unnecessary.
@@ -856,7 +858,7 @@ Reads and writes hang on Named Pipes until there is at least one reader and one
Hello
\end{lstlisting}
Any time \keyword{open} is called on a named pipe, the kernel blocks until another process calls the opposite open. Meaning, echo calls \keyword{open(..,\ O\_RDONLY)} but that blocks until cat calls \keyword{open(..,\ O\_WRONLY)}, then the programs are allowed to continue.
Any time \keyword{open} is called on a named pipe, the kernel blocks until another process calls the opposite open. Meaning, echo calls \keyword{open(..,\ O\_WRONLY)} but that blocks until cat calls \keyword{open(..,\ O\_RDONLY)}, then the programs are allowed to continue.
\subsection{Race condition with named pipes}
@@ -895,7 +897,7 @@ Sometimes it looks like it works because the execution of the code looks somethi
Time 1 & open(O\_RDWR) \& write() & \\
Time 2 & & open(O\_RDONLY) \& read() \\
Time 3 & close() \& exit() & \\
Time 4 & print() \& exit() & \\
Time 4 & & print() \& exit() \\
\end{tabular}
\end{table}
\end{center}
+14 -10
View File
@@ -234,12 +234,12 @@ We don't even have to look through the entire heap!
\caption{First fit finds the first match}
\end{figure}
One thing to keep in mind is that those placement strategies don't need to replace the block.
One thing to keep in mind is that those placement strategies don't need to split the block.
For example, our first fit allocator could've returned the original block unbroken.
Notice that this would lead to about 14KiB of space being unused by the user and the allocator.
We call this internal fragmentation.
In contrast, external fragmentation is that even though we have enough memory in the heap, it may be divided up in a way so a continuous block of that size is unavailable.
In contrast, external fragmentation is that even though we have enough memory in the heap, it may be divided up in a way so a contiguous block of that size is unavailable.
In our previous example, of the 64KiB of heap memory, 17KiB is allocated, and 47KiB is free.
However, the largest available block is only 30KiB because our available unallocated heap memory is fragmented into smaller pieces.
@@ -263,7 +263,7 @@ For those who are interested, the paper concludes that actual memory usage over
The problem with this analysis is that few real-world applications need this type of one-shot allocation.
Video game object allocations will typically designate a different subheap for each level and fill up that subheap if they need a quick memory allocation scheme that they can throw away.
In practice, we'll be using the result from a more rigorous survey conducted in 2005 \cite{10.1007/3-540-60368-9_19}.
In practice, we'll be using the result from a more rigorous survey conducted in 1995 \cite{10.1007/3-540-60368-9_19}.
The survey makes sure to note that memory allocation is a moving target.
A good allocation scheme to one program may not be a good allocation scheme for another program.
Programs don't uniformly follow the distribution of allocations.
@@ -290,7 +290,7 @@ Some additional notes we make
Using Fibonacci heaps, however, could be extremely inefficient.
\item First-Fit needs to have a block order. Most of the time programmers will default to linked lists which is a fine choice. There aren't too many improvements you can make with a least recently used and most recently used linked list policy, but with address ordered linked lists you can speed up insertion from O(n) to O(log(n)) by using a randomized skip-list in conjunction with your singly-linked list.
An insert would use the skip list as shortcuts to find the right place to insert the block and removal would go through the list as normal.
\item There are many placement strategies that we haven't talked about, one is next-fit which is first fit on the next fit block. This adds deterministic randomness -- pardon the oxymoron. You won't be expected to know this algorithm, but know that as you are implementing a memory allocator as part of a machine problem, there are more than these.
\item There are many placement strategies that we haven't talked about, one is next-fit, which is like first fit except that each search resumes where the previous search stopped instead of starting at the beginning of the heap. This adds deterministic randomness -- pardon the oxymoron. You won't be expected to know this algorithm, but know that as you are implementing a memory allocator as part of a machine problem, there are more than these.
\end{enumerate}
\section{Memory Allocator Tutorial}
@@ -337,15 +337,19 @@ typedef struct {
size_t block_size;
char data[0];
} block;
// Stored at the end of each block (bt in the figures above)
typedef struct {
size_t block_size;
} boundary_tag;
block *p = sbrk(100);
p->size = 100 - sizeof(*p) - sizeof(BTag);
p->block_size = 100 - sizeof(*p) - sizeof(boundary_tag);
// Other block allocations
\end{lstlisting}
We could navigate from one block to the next block by adding the block's size.
\begin{lstlisting}[language=C]
p + sizeof(metadata) + p->block_size + sizeof(BTag)
p + sizeof(metadata) + p->block_size + sizeof(boundary_tag)
\end{lstlisting}
Make sure to get your casting right!
@@ -382,7 +386,7 @@ typedef struct {
char data[0];
} block;
block *p = sbrk(100);
p->size = 100 - sizeof(*p) - sizeof(boundary_tag);
p->block_size = 100 - sizeof(*p) - sizeof(boundary_tag);
// Other block allocations
\end{lstlisting}
@@ -451,10 +455,10 @@ There is a \emph{lot} of overhead for that allocation which is what we are tryin
When \keyword{free} is called we need to re-apply the offset to get back to the `real' start of the block -- to where we stored the size information.
A naive implementation would simply mark the block as unused.
If we are storing the block allocation status in a bitfield, then we need to clear the bit:
If we are storing the block allocation status in a bitfield, then we need to set the \keyword{is\_free} bit:
\begin{lstlisting}[language=C]
p->info.is_free = 0;
p->info.is_free = 1;
\end{lstlisting}
However, we have a bit more work to do.
@@ -643,7 +647,7 @@ See \href{http://man7.org/linux/man-pages/man3/malloc.3.html}{the man page} or t
\item
What is the First Fit Placement strategy? It's a little bit better with Fragmentation, right? Expected Time Complexity?
\item
Let's say that we are using a buddy allocator with a new slab of 64kb. How does it go about allocating 1.5kb?
Let's say that we are using a buddy allocator with a new slab of 64KiB. How does it go about allocating 1.5KiB?
\item
When does the 5 line \keyword{sbrk} implementation of malloc have a use?
\item
+21
View File
@@ -59,3 +59,24 @@
month = aug,
abstract = {},
}
@misc{google_ipv6_stats,
title = {{IPv6 Adoption Statistics}},
author = {{Google}},
url = {https://www.google.com/intl/en/ipv6/statistics.html},
year = 2026,
note = {Accessed September 2026},
}
@misc{rfc9110,
series = {Request for Comments},
number = 9110,
howpublished = {RFC 9110},
publisher = {RFC Editor},
doi = {10.17487/RFC9110},
url = {https://www.rfc-editor.org/rfc/rfc9110},
author = {Roy T. Fielding and Mark Nottingham and Julian Reschke},
title = {{HTTP Semantics}},
year = 2022,
month = jun,
}
+18 -15
View File
@@ -61,7 +61,7 @@ As for another definition, a protocol is a set of specifications put forward by
The following is a short introduction to internet protocol (IP), the primary way to send datagrams of information from one machine to another.
``IP4'', or more precisely, IPv4 is version 4 of the Internet Protocol that describes how to send packets of information across a network from one machine to another.
Even as of 2018, IPv4 still dominates Internet traffic, but Google reports that 24 countries now supply 15\% of their traffic through IPv6 \cite{internet_society_2018}.
IPv4 carried most Internet traffic for decades and is still widely used today, although IPv6 is catching up (see below).
A significant limitation of IPv4 is that source and destination addresses are limited to 32 bits.
IPv4 was designed at a time when the idea of 4 billion devices connected to the same network was unthinkable or at least not worth making the packet size larger.
IPv4 addresses are written typically in a sequence of four octets delimited by periods "255.255.255.0" for example.
@@ -70,7 +70,8 @@ Each IPv4 datagram includes a small header - typically 20 octets, that includes
Conceptually the source and destination addresses can be split into two: a network number the upper bits and lower bits represent a particular host number on that network.
A newer packet protocol IPv6 solves many of the limitations of IPv4 like making routing tables simpler and 128-bit addresses.
However, little web traffic is IPv6 based on comparison as of 2018 \cite{internet_society_2018}.
Adoption was slow at first: in 2018, only a small fraction of web traffic used IPv6 \cite{internet_society_2018}.
By the mid-2020s, Google measured roughly half of its user traffic arriving over IPv6 \cite{google_ipv6_stats}.
We write IPv6 addresses in a sequence of eight, four hexadecimal delimiters like "1F45:0000:0000:0000:0000:0000:0000:0000".
Since that can get unruly, we can omit the zeros "1F45::". A machine can have an IPv6 address and an IPv4 address.
@@ -89,9 +90,8 @@ This book covers how IP deals with routing, fragmenting, and reassembling upper-
\end{figure}
One of the big features of IPv6 is the address space.
The world ran out of IP addresses a while ago and has been using hacks to get around that.
The world ran out of IPv4 addresses years ago (IANA handed out its last free blocks in 2011) and has been using hacks to get around that.
With IPv6 there are enough internal and external addresses so even if we discover alien civilizations, we probably won't run out.
The other benefit is that these addresses are leased not bought, meaning that if something drastic happens in let's say the Internet of things and there needs to be a change in the block addressing scheme, it can be done.
Another big feature is security through IPsec.
IPv4 was designed with little to no security in mind.
@@ -331,7 +331,7 @@ For network communications, we need to standardize on the agreed format.
Any longer integers need to have the computers specify the order.
These functions are read as `host to network'.
The inverse functions (\keyword{ntohs}, \keyword{ntohl}) convert network ordered byte values to host-ordered ordering.
The inverse functions (\keyword{ntohs}, \keyword{ntohl}) convert network ordered byte values to host byte ordering.
So, is host-ordering little-endian or big-endian?
The answer is - it depends on your machine!
It depends on the actual architecture of the host running the code.
@@ -571,7 +571,7 @@ If the client had requested a non-existent path, e.g.
HTTP/1.1 404 Not Found
\end{lstlisting}
For more information, RFC 7231 has the most current specifications on the most common HTTP method today \cite{rfc7231}.
For more information, RFC 9110 (which replaced RFC 7231 in 2022) is the current specification of HTTP semantics, including the request methods and status codes \cite{rfc9110}.
\section{Layer 4: TCP Server}
@@ -1029,7 +1029,7 @@ For example, a coffee shop internet connection could easily subvert your DNS req
The way this is usually subverted is that after the IP address is obtained then a connection is usually made over HTTPS.
HTTPS uses what is called the TLS (formerly known as SSL) to secure transmissions and verify that the hostname is recognized by a Certificate Authority.
Certificate Authorities often get hacked so be careful of equating a green lock to secure.
Even with this added layer of security, the United States government has recently issued a request for everyone to upgrade their DNS to DNSSec which includes additional security-focused technologies to verify with high probability that an IP address is truly associated with a hostname.
Even with this added layer of security, the United States government has required its federal agencies to deploy DNSSEC since 2008; DNSSEC includes additional security-focused technologies to verify with high probability that an IP address is truly associated with a hostname.
Digression aside, DNS works like this in a nutshell:
\begin{enumerate}
@@ -1285,10 +1285,10 @@ Some of the more common ones will be detailed below.
There are several problems with using epoll. Here we will detail a few.
\begin{enumerate}
\item There are two modes. Level triggered and edge-triggered. Level triggered means that while the file descriptor has events on it, it will be returned by epoll when calling the ctl function. In edge-triggered, the caller will only get the file descriptor once it goes from zero events to an event.
\item There are two modes. Level triggered and edge-triggered. Level triggered means that while the file descriptor has events on it, it will be returned by every call to \keyword{epoll\_wait}. In edge-triggered, the caller will only get the file descriptor once it goes from zero events to an event.
This means if you forget to read, write, accept etc on the file descriptor until you get an EWOULDBLOCK, that file descriptor will be dropped.
\item If at any point you duplicate a file descriptor and add it to epoll, you will get an event from that file descriptor and the duplicated one.
\item You can add an epoll object to epoll. Edge triggered and level-triggered modes are the same because ctl will reset the state to zero events.
\item You can add an epoll object to epoll.
\item Depending on the conditions, you may get a file descriptor that was closed from Epoll. This isn't a bug. The reason that this happens is epoll works on the kernel object level, not the file descriptor level.
If the kernel object lives longer and the right flags are set, a process could get a closed file descriptor.
This also means that if you close the file descriptor, there is no way to remove the kernel object.
@@ -1322,25 +1322,28 @@ The stub code is the necessary code to hide the complexity of performing a remot
One of the roles of the stub code is to \emph{marshal} the necessary data into a format that can be sent as a byte stream to a remote server.
\begin{lstlisting}[language=C]
// On the outside, 'getHiscore' looks like a normal function call
// On the outside, 'getHighScore' looks like a normal function call
// On the inside, the stub code performs all of the work to send and receive data to and from the remote machine.
int getHighScore(char* game) {
// Marshal the request into a sequence of bytes:
char* buffer;
asprintf(&buffer,"getHiscore(%s)!", name);
asprintf(&buffer,"getHighScore(%s)!", game);
// Send down the wire (we do not send the zero byte; the '!' signifies the end of the message)
// fd is an already-connected socket to the remote server
write(fd, buffer, strlen(buffer) );
free(buffer);
// Wait for the server to send a response
ssize_t bytesread = read(fd, buffer, sizeof(buffer));
char response[64];
ssize_t bytesread = read(fd, response, sizeof(response) - 1);
if (bytesread < 0) return -1;
// Example: unmarshal the bytes received back from text into an int
buffer[bytesread] = 0; // Turn the result into a C string
response[bytesread] = 0; // Turn the result into a C string
int score= atoi(buffer);
free(buffer);
int score = atoi(response);
return score;
}
\end{lstlisting}
+2 -2
View File
@@ -2,7 +2,7 @@
\epigraph{Hindsight is 20-20}{Unknown}
This chapter is meant to serve as a big "why are we learning all of this".
This chapter is meant to serve as a big ``why are we learning all of this''.
In all of your previous classes, you were learning what to do.
How to program a data structure, how to code a for loop, how to prove something.
This is the first class that is largely focused on what \textit{not} to do.
@@ -161,7 +161,7 @@ Lessons Learned: Get an antivirus and/or apparmor and make sure that an applicat
\section{The Woes of Shell Scripting}
Required Sections: Intro to C/Appending
Required Sections: Appendix/Shell
\href{https://www.pcworld.com/article/2871653/scary-steam-for-linux-bug-erases-all-the-personal-files-on-your-pc.html}{Steam}
+8 -8
View File
@@ -480,7 +480,7 @@ The following snippet contains only one description.
\begin{lstlisting}[language=C]
int file = open(...);
if(!fork) {
if(!fork()) {
read(file, ...);
} else {
read(file, ...);
@@ -491,7 +491,7 @@ One process will read one part of the file, the other process will read another
In the following example, there are two descriptions caused by two different file handles.
\begin{lstlisting}[language=C]
if(!fork) {
if(!fork()) {
int file = open(...);
read(file, ...);
} else {
@@ -516,8 +516,7 @@ size_t buffer_cap = 0;
char * buffer = NULL;
ssize_t nread;
FILE * file = fopen("test.txt", "r");
int count = 0;
while((nread = getline(&buffer, &buffer_cap, file) != -1) {
while((nread = getline(&buffer, &buffer_cap, file)) != -1) {
printf("%s", buffer);
if(fork() == 0) {
exit(0);
@@ -546,8 +545,7 @@ size_t buffer_cap = 0;
char * buffer = NULL;
ssize_t nread;
FILE * file = fopen("test.txt", "r");
int count = 0;
while((nread = getline(&buffer, &buffer_cap, file) != -1) {
while((nread = getline(&buffer, &buffer_cap, file)) != -1) {
printf("%s", buffer);
fflush(file);
if(fork() == 0) {
@@ -581,6 +579,7 @@ waitpid(child, NULL, 0);
If you are interested in how this works, check out the appendix for a description of the Fork-file problem.
\section{Waiting and Executing}
\label{sec:waiting_and_executing}
If the parent process wants to wait for the child to finish, it must use \keyword{waitpid} (or \keyword{wait}), both of which wait for a child to change process states, which can be one of the following:
@@ -809,7 +808,7 @@ int main() {
}
\end{lstlisting}
The example writes "Captain's Log" to a file then prints everything in /usr/include to the same file.
The example writes "Captain's log" to a file then prints everything in /usr/include to the same file.
There's no error checking in the above code (we assume close, open, chdir etc. work as expected).
\begin{enumerate}
@@ -926,7 +925,8 @@ setenv("HOME", "/home/user", 1 /*set overwrite to true*/ );
\end{lstlisting}
Environment variables are important because they are inherited between processes and can be used to specify a standard set of behaviors \cite{env_std_2018}, although you don't need to memorize the options.
Another security related concern is that environment variables cannot be read by an outside process, whereas argv can be.
Another security related concern is that argv is visible to other users, for example in the output of \keyword{ps}, whereas environment variables are not displayed there.
However, environment variables are not secret: on Linux, a process running as the same user (or as root) can read them from \keyword{/proc/<pid>/environ}.
\section{Further Reading}
+9 -5
View File
@@ -147,7 +147,9 @@ char* toString(int a, char*mesg, double val, void* ptr) {
\subsection{Input parsing}
\begin{enumerate}
\item Why should you check the return value of sscanf and scanf? \#\# Q 5.2 Why is `gets' dangerous?
\item Why should you check the return value of sscanf and scanf?
\item Why is \keyword{gets} dangerous?
\item Write a complete program that uses \keyword{getline}. Ensure your program has no memory leaks.
@@ -155,8 +157,10 @@ char* toString(int a, char*mesg, double val, void* ptr) {
\item What mistake did the programmer make in the following code? Is it possible to fix it?
i) using heap memory?
ii) using global (static) memory?
\begin{enumerate}
\item using heap memory?
\item using global (static) memory?
\end{enumerate}
\begin{lstlisting}[language=C]
static int id;
@@ -563,7 +567,7 @@ $ chmod 000 secret.txt
\item Which common protocol is stream-based and will resend data if packets are lost?
\item What is the SYN ACK ACK-SYN handshake?
\item What is the SYN, SYN-ACK, ACK handshake?
\item Which one of the following is NOT a feature of TCP?
\begin{enumerate}
@@ -604,7 +608,7 @@ $ chmod 000 secret.txt
\item How does Network Address Translation (NAT) work?
\item Assuming a network has a 20ms One Way Transit Time between Client and Server, how much time would it take to establish a TCP Connection?
\item Assuming a network has a 20ms One Way Transit Time between Client and Server, how much time does the TCP three-way handshake (SYN, SYN-ACK, ACK) take to complete, measured from when the client sends the SYN until the server receives the final ACK?
\begin{enumerate}
\item 20ms
\item 40ms
+7 -6
View File
@@ -186,7 +186,7 @@ For this exposition, we will simplify this discussion to use the total running t
\subsection{Preemptive Shortest Job First (PSJF)}
Preemptive shortest job first is like shortest job first but if a new job comes in with a shorter runtime than the total runtime of the current job, it is run instead.
If it is equal like our example our algorithm can choose.
If the runtimes are equal, our algorithm can choose either.
The scheduler uses the \emph{total} runtime of the process.
If the scheduler wants to compare the shortest \emph{remaining} time left, that is a variant of PSJF called Shortest Remaining Time First (SRTF).
@@ -212,12 +212,13 @@ If the scheduler wants to compare the shortest \emph{remaining} time left, that
Here's what our algorithm does.
It runs P2 because it is the only thing to run.
Then P1 comes in at 1000ms, P2 runs for 2000ms, so our scheduler preemptively stops P2, and lets P1 run all the way through.
This is completely up to the algorithm because the times are equal.
Then P1 comes in at 1000ms, P2 runs for 2000ms, so our scheduler preemptively stops P2, and lets P1 run all the way through, because P1's total runtime (1000ms) is shorter than P2's (2000ms).
P2 then runs to completion.
Then, P5 comes in -- since no processes are running, the scheduler will run process 5.
P4 comes in, and since the runtimes are equal to P5, the scheduler stops P5 and runs P4.
P4 comes in, and since its total runtime (4000ms) is shorter than P5's (5000ms), the scheduler stops P5 and runs P4.
Finally, P3 comes in, preempts P4, and runs to completion.
Then P4 runs, then P5 runs.
Under SRTF (shortest remaining time first), which compares remaining rather than total time, each of these preemptions (when P1, P4 and P3 arrive) would be a tie.
\textbf{Advantages}
@@ -281,8 +282,8 @@ You can see the convoy effect for P5.
Processes are scheduled in order of their arrival in the ready queue.
After a small time step though, a running process will be forcibly removed from the running state and placed back on the ready queue.
This ensures long-running processes refrain from starving all other processes from running.
The maximum amount of time that a process can execute before being returned to the ready queue is called the time quanta.
As the time quanta approaches infinity, Round Robin will be equivalent to FCFS.
The maximum amount of time that a process can execute before being returned to the ready queue is called the time quantum.
As the time quantum approaches infinity, Round Robin will be equivalent to FCFS.
\begin{figure}[H]
\centering
+16 -13
View File
@@ -179,14 +179,14 @@ The following snippet is a high-level proof of concept.
\begin{lstlisting}[language=C]
char *a[10];
for (int i = 10; i != 1; --i) {
a[i] = calloc(1, 1);
for (int k = 9; k != 0; --k) {
a[k] = calloc(1, 1);
}
a[0] = 0xCAFE;
a[0] = (char *) 0xCAFE;
int val;
int j = 10; // This will be in a register
int i = 10; // This will be in main memory
for (int i = 10; i != 0; --i, --j) {
int j = 9; // This will be in a register
int i; // This will be in main memory
for (i = 9; i >= 0; --i, --j) {
if (i) {
val = *a[j];
}
@@ -194,8 +194,9 @@ for (int i = 10; i != 0; --i, --j) {
\end{lstlisting}
Let's analyze this code.
The first loop allocates 9 elements through a valid malloc.
The last element is \keyword{0xCAFE}, meaning a dereference should result in a SEGFAULT.
The first loop allocates 9 elements, \keyword{a[1]} through \keyword{a[9]}, through a valid calloc.
The last element that the second loop visits, \keyword{a[0]}, is \keyword{0xCAFE}, meaning a dereference should result in a SEGFAULT.
The second loop runs 10 times.
For the first 9 iterations, the branch is taken and \keyword{val} is assigned to a valid value.
The interesting part happens in the last iteration.
The resulting behavior of the program is to skip the last iteration.
@@ -317,7 +318,7 @@ As more and more of our systems are hacked over the web, it is important to unde
\begin{enumerate}
\item Encryption.
\textbf{TCP is unencrypted!} This means any data that is sent over a TCP connection is in plain text.
If one needs to send encrypted data, one needs to use a higher level protocol such as HTTPs or develop their own.
If one needs to send encrypted data, one needs to use a higher level protocol such as HTTPS or develop their own.
\item Identity Verification.
In TCP, there is no way to verify the identity of who the program is connecting to.
There are no checks or federated databases in place.
@@ -341,11 +342,13 @@ As more and more of our systems are hacked over the web, it is important to unde
\subsection{Security at the DNS Level}
As of 2019, the United States Department of Homeland Security released a directive to switch all services from DNS to DNSSec \url{https://www.cisa.gov/news-events/directives/ed-19-01-mitigate-dns-infrastructure-tampering-closed}.
This directive is an inherent flaw of the DNS system.
First, DNS doesn't offer any sort of verification on domain name requests.
In January 2019, the United States Department of Homeland Security issued Emergency Directive 19-01, which required federal agencies to audit and lock down their DNS records after a wave of DNS infrastructure tampering \url{https://www.cisa.gov/news-events/directives/ed-19-01-mitigate-dns-infrastructure-tampering-closed}.
Attacks like these exploit weaknesses built into DNS itself.
DNSSEC is an extension to DNS, not a replacement for it, that lets a resolver verify that a response was signed by the owner of the domain.
First, plain DNS doesn't offer any sort of verification on domain name requests.
That is, it is easy to spoof DNS nameservers such that they point your browser to potentially malicious servers.
Remember that DNS requests are sent as unsecured UDP packets, which are prone to tampering. This means that if an attacker snags a plain-text request for a DNS server, that attacker can now send the result back to the requester.
Remember that traditional DNS requests are sent as unsecured UDP packets, which are prone to tampering. This means that if an attacker snags a plain-text request for a DNS server, that attacker can now send the result back to the requester.
Many operating systems and browsers can now send DNS requests over an encrypted connection instead, using DNS-over-TLS (DoT) or DNS-over-HTTPS (DoH), which protects them from being read or altered on the way to the resolver.
More commonly instead of just attacking one person, they will connect to a public wifi station and poison the cache of the router -- meaning that all who are connected will get a bad IP address when requesting a domain name.
This can get into serious spoofing attacks if one tries to pretend they are a major bank.
+15 -15
View File
@@ -13,8 +13,8 @@ For those of you with an architecture background, the interrupts used here aren'
Those interrupts are almost always handled by the kernel because they require higher levels of privileges.
Instead, we are talking about software interrupts that are generated by the kernel -- though they can be in response to a hardware event like SIGSEGV.
This chapter will go over how to read information from a process that has either exited or been signaled.
Then, it will deep dive into what signals are, how the kernel deals with a signal, and the various ways processes can handle signals both with and without threads.
Reading the exit status of a process that has either exited or been signaled (\keyword{WIFSIGNALED}, \keyword{WTERMSIG}, etc.) is covered in Section~\ref{sec:waiting_and_executing} of the Processes chapter.
This chapter will deep dive into what signals are, how the kernel deals with a signal, and the various ways processes can handle signals both with and without threads.
\section{The Deep Dive of Signals}
@@ -183,19 +183,19 @@ $ kill -SIGINT 4409
$ kill -SIGKILL 4409
$ kill -9 4409
# Use kill all instead to kill a process by executable name
$ killall -l firefox
$ killall firefox
\end{lstlisting}
To send a signal to the running process, use \keyword{raise} or \keyword{kill} with \keyword{getpid()}.
\begin{lstlisting}[language=C]
raise(int sig); // Send a signal to myself!
kill(getpid(), int sig); // Same as above
raise(sig); // Send a signal to myself!
kill(getpid(), sig); // Same as above
\end{lstlisting}
For non-root processes, signals can only be sent to processes of the same user.
You can't SIGKILL any process!
\keyword{man -s2 kill} for more details.
You can't SIGKILL just any process!
\keyword{man 2 kill} for more details.
\section{Handling Signals}
@@ -205,16 +205,16 @@ Re-entrant safety means that your function can be frozen at any point and execut
Let's take the following
\begin{lstlisting}[language=C]
int func(const char *str) {
void func(const char *str) {
static char buffer[200];
strncpy(buffer, str, 199);
# Here is where we get paused
printf("%s\n", buffer)
// Here is where we get paused
printf("%s\n", buffer);
}
\end{lstlisting}
\begin{enumerate}
\item We execute \keyword(func("Hello"))
\item We execute \keyword{func("Hello")}
\item The string gets copied over to the buffer completely (strcmp(buffer, "Hello") == 0)
\item A signal is delivered and the function state freezes, we also stop accepting any new signals until after the handler (we do this for convenience)
\item We execute \keyword{func("World")}
@@ -251,7 +251,7 @@ int main() {
The above code might appear to be correct on paper.
However, we need to provide a hint to the compiler and the CPU core that will execute the \keyword{main()} loop.
We need to prevent compiler optimization.
The expression \keyword{\!pleaseStop} doesn't get changed in the body of the loop, so some compilers will optimize it to \keyword{true} \todo{citation needed}.
The expression \keyword{!pleaseStop} doesn't get changed in the body of the loop, so some compilers will optimize it to \keyword{true} \todo{citation needed}.
Secondly, we need to ensure that the value of \keyword{pleaseStop} is uncached using a CPU register and instead always read from and written to main memory.
The \keyword{sig\_atomic\_t} type implies that all the bits of the variable can be read or modified as an \keyword{atomic operation} - a single uninterruptible operation.
It is impossible to read a value that is composed of some new bit values and old bit values.
@@ -307,7 +307,7 @@ struct sigaction sa;
sa.sa_handler = myhandler;
sigemptyset(&sa.sa_mask);
sa.sa_flags = 0;
sigaction(SIGALRM, &sa, NULL)
sigaction(SIGALRM, &sa, NULL);
\end{lstlisting}
However, we typically may also set the mask and the flags field.
@@ -350,7 +350,7 @@ It is a common error to forget to initialize the signal set before adding to the
\begin{lstlisting}[language=C]
sigset_t set, oldset;
sigaddset(&set, SIGINT); // Ooops!
sigprocmask(SIG_SETMASK, &set, &oldset)
sigprocmask(SIG_SETMASK, &set, &oldset);
\end{lstlisting}
Correct code initializes the set to be all on or all off. For example,
@@ -463,7 +463,7 @@ pthread_sigmask(SIG_BLOCK, &set, NULL) - add the signal set to the thread's mask
pthread_sigmask(SIG_UNBLOCK, &set, NULL) - remove the signal set from the thread's mask
\end{lstlisting}
A signal then can be delivered to any signal thread that is willing to accept that signal.
A signal then can be delivered to any thread that is willing to accept that signal.
If two or more threads can receive the signal then which thread will be interrupted is arbitrary!
A common practice is to have one thread that can receive all signals or if there is a certain signal that requires special logic, have multiple threads for multiple signals.
Even though programs from the outside can't send signals to specific threads, you can do that internally with \keyword{pthread\_kill(pthread\_t thread, int sig)}.
+31 -27
View File
@@ -41,7 +41,7 @@ int main() {
}
\end{lstlisting}
A typical output of the above code is \keyword{ARGGGH sum is <some number less than expected>} because there is a race condition.
A typical output of the above code is \keyword{ARRRRG sum is <some number less than expected>} because there is a race condition.
The code allows two threads to read and write \keyword{sum} at the same time.
For example, both threads copy the current value of sum into CPU that runs each thread (let's pick 123).
Both threads increment one to their own copy.
@@ -203,9 +203,9 @@ return NULL;
}
\end{lstlisting}
This process runs slower because we lock and unlock the mutex a million times, which is expensive - at least compared with incrementing a variable.
In this simple example, we didn't need threads - we could have added up twice!
A faster multi-thread example would be to add one million using an automatic (local) variable and only then adding it to a shared total after the calculation loop has finished:
This process runs slower because we lock and unlock the mutex ten million times, which is expensive - at least compared with incrementing a variable.
In this simple example, we didn't need threads.
A faster multi-thread example would be to add ten million using an automatic (local) variable and only then adding it to a shared total after the calculation loop has finished:
\begin{lstlisting}[language=C]
int local = 0;
@@ -234,7 +234,7 @@ The following code creates a mutex that does effectively nothing.
\begin{lstlisting}[language=C]
int a;
pthread_mutex_t m1 = PTHREAD_MUTEX_INITIALIZER,
m2 = = PTHREAD_MUTEX_INITIALIZER;
m2 = PTHREAD_MUTEX_INITIALIZER;
// later
// Thread 1
pthread_mutex_lock(&m1);
@@ -302,7 +302,7 @@ void unlock(mutex_t *m) {
Version 1 uses `busy-waiting' unnecessarily wasting CPU resources.
However, there is a more serious problem.
We have a race-condition!
If two threads both called \keyword{lock} concurrently, it is possible that both threads would read \keyword{m\_locked} as zero.
If two threads both called \keyword{lock} concurrently, it is possible that both threads would read \keyword{m->locked} as zero.
Thus both threads would believe they have exclusive access to the lock and both threads will continue.
We might attempt to reduce the CPU overhead a little by calling \keyword{pthread\_yield()} inside the loop - pthread\_yield suggests to the operating system that the thread does not use the CPU for a short while, so the CPU may be assigned to threads that are waiting to run.
@@ -345,7 +345,7 @@ int mutex_init(mutex* mtx){
\end{lstlisting}
This is the initialization code, nothing fancy here.
We set the state of the mutex to unlocked and set the owner to locked.
We set the state of the mutex to unlocked and set the owner to unassigned.
\begin{lstlisting}[language=C]
int mutex_lock(mutex* mtx){
@@ -776,7 +776,7 @@ int main() {
Before we fix the problems with semaphores,
how would we fix the problems with condition variables?
Try it out before you look at the code in the previous section.
Try it out before you look at the code below.
We need to wait in push and pop if our stack is full or empty respectively.
Attempted solution:
@@ -800,7 +800,7 @@ double pop(stack_t *s) {
}
\end{lstlisting}
Does the following solution work?
Does the above solution work?
Take a second before looking at the answer to spot the errors.
So did you catch all of them?
@@ -860,12 +860,13 @@ double pop() {
// Wait until there's at least one item
sem_wait(&sitems);
...
}
void push(double v) {
// Wait until there's at least one space
sem_wait(&sremain);
...
}
void push(double v) {
// Wait until there's at least one space
sem_wait(&sremain);
...
}
\end{lstlisting}
Sketch \#2 has implemented the \keyword{post} too early.
@@ -926,7 +927,7 @@ sem_t sitems, sremain;
void init() {
sem_init(&sitems, 0, 0);
sem_init(&sremains, 0, SPACES); // 10 spaces
sem_init(&sremain, 0, SPACES); // 10 spaces
}
double pop() {
@@ -1191,8 +1192,8 @@ lower my flag
\end{lstlisting}
This solution satisfies Mutual Exclusion, Bounded Wait and Progress.
If thread \#2 has set turn to 2 and is currently inside the critical section.
Thread \#1 arrives, \emph{sets the turn back to 1} and now waits until thread 2 lowers the flag.
Suppose thread \#2 has set turn to 1 and is currently inside the critical section.
Thread \#1 arrives, \emph{sets the turn to 2} and now waits until thread 2 lowers the flag.
\begin{enumerate}
\item Mutual Exclusion. Let's try to sketch a simple proof again.
@@ -1225,7 +1226,7 @@ We don't want to call malloc while implementing a primitive, or we may deadlock!
typedef struct sem_t {
ssize_t count;
pthread_mutex_t m;
pthread_condition_t cv;
pthread_cond_t cv;
} sem_t;
\end{lstlisting}
\end{itemize}
@@ -1598,7 +1599,8 @@ Which brings us to attempt \#3.
\subsection{Attempt \#3}
Remember that \keyword{pthread\_cond\_wait} performs \emph{Three} actions.
Firstly, it atomically unlocks the mutex and then sleeps (until it is woken by \keyword{pthread\_cond\_signal} or \keyword{pthread\_cond\_broadcast}).
Firstly, it unlocks the mutex.
Secondly, it sleeps (until it is woken by \keyword{pthread\_cond\_signal} or \keyword{pthread\_cond\_broadcast}); these first two actions happen atomically.
Thirdly, the awoken thread must re-acquire the mutex lock before returning.
Thus only one thread can actually be running inside the critical section defined by the lock and unlock() methods.
@@ -2084,27 +2086,29 @@ See \href{http://stackoverflow.com/questions/19172541/procs-fork-and-mutexes}{st
So what should we do? We should use a shared mutex! Consider the following code.
\begin{lstlisting}[language=C]
pthread_mutex_t * mutex = NULL;
pthread_mutexattr_t attr;
pthread_mutex_t * pmutex = NULL;
pthread_mutexattr_t attrmutex;
void write_string(const char *data) {
pthread_mutex_lock(mutex);
pthread_mutex_lock(pmutex);
int fd = open("my_file.txt", O_WRONLY);
int bytes_to_write = strlen(data), written = 0;
while(written < bytes_to_write) {
written += write(fd, data + written, bytes_to_write - written);
ssize_t result = write(fd, data + written, bytes_to_write - written);
if(result == -1) break; // give up on an error
written += result;
}
close(fd);
pthread_mutex_unlock(mutex);
pthread_mutex_unlock(pmutex);
}
int main() {
pthread_mutexattr_init(&attr);
pthread_mutexattr_setpshared(&attr, PTHREAD_PROCESS_SHARED);
pthread_mutexattr_init(&attrmutex);
pthread_mutexattr_setpshared(&attrmutex, PTHREAD_PROCESS_SHARED);
pmutex = mmap (NULL, sizeof(pthread_mutex_t),
PROT_READ|PROT_WRITE, MAP_SHARED|MAP_ANON, -1, 0);
pthread_mutex_init(pmutex, &attrmutex);
if(!fork()) {
if(fork()) { // parent
write_string("key1: value1");
wait(NULL);
pthread_mutex_destroy(pmutex);
+15 -15
View File
@@ -29,7 +29,7 @@ On the other hand, creating threads is more useful when:
\begin{itemize}
\item You want to leverage the power of a multi-core system to do one task
\item When you can't deal with the overhead of processes
\item When you want communication between the processes simplified
\item When you want communication between the threads simplified
\item When you want threads to be part of the same process
\end{itemize}
@@ -41,7 +41,7 @@ If the thread calls another function, we move our stack pointer down, so that we
Once it returns from a function, we can move the stack pointer back up to its previous value.
We keep a copy of the old stack pointer value - on the stack!
This is why returning from a function is quick.
It's easy to `free' the memory used by automatic variables because the program needs to change the stack pointer.
It's easy to `free' the memory used by automatic variables because the program only needs to change the stack pointer.
In a multi-threaded program, there are multiple stacks but only one address space. The pthread library allocates some stack space and uses the \keyword{clone} function call to start the thread at that stack address.
@@ -321,7 +321,7 @@ int main() {
}
\end{lstlisting}
Race conditions aren't in our code.
Race conditions aren't only in our code.
They can be in provided code.
Some functions like \keyword{asctime}, \keyword{getenv}, \keyword{strtok}, \keyword{strerror} are not thread-safe.
Let's look at a simple function that is also not `thread-safe'.
@@ -346,15 +346,14 @@ Here is one valid solution.
\begin{lstlisting}[language=C]
int to_message_r(int num, char *buf, size_t nbytes) {
size_t written;
int written;
if (num < 10) {
written = snprintf(buf, nbtytes, "%d : blah blah" , num);
written = snprintf(buf, nbytes, "%d : blah blah" , num);
} else {
strncpy(buf, "Unknown", nbytes);
buf[nbytes] = '\0';
written = strlen(buf) + 1;
written = snprintf(buf, nbytes, "%s", "Unknown");
}
return written <= nbytes;
// Nonzero if the whole message (and its '\0') fit in buf
return written >= 0 && (size_t) written < nbytes;
}
\end{lstlisting}
@@ -434,7 +433,8 @@ void merge_sort(int *arr, size_t len){
With your new understanding of threads, all you need to do is create a thread for the left half, and one for the right half.
Given that your CPU has multiple real cores, you will see a speedup following \href{https://en.wikipedia.org/wiki/Amdahl's_law}{Amdahl's Law}.
The time complexity analysis gets interesting here as well.
The parallel algorithm runs in $O(\log^3(n))$ running time because the analysis assumes that we have a lot of cores.
Even if we assume that we have as many cores as we need, the algorithm as written runs in $O(n)$ time, because the final merge of the two halves is still done sequentially by a single thread.
Parallelizing the merge step as well brings the running time down to $O(\log^2(n))$.
In practice though, we typically do two changes.
One, once the array gets small enough, we ditch the Parallel Merge Sort algorithm and do a conventional sort that works fast on small arrays, usually cache coherency rules at this level.
@@ -492,13 +492,13 @@ Take a look at the example code below.
\begin{lstlisting}[language=C]
// 8 KiB stacks
// 8 MiB stacks
#define STACK_SIZE (8 * 1024 * 1024)
int thread_start(void *arg) {
// Just like the pthread function
puts("Hello Clone!")
// This share the same heap and address space!
puts("Hello Clone!");
// This shares the same heap and address space!
return 0;
}
@@ -507,10 +507,10 @@ int main() {
char *child_stack = malloc(STACK_SIZE);
// Remember stacks work by growing down, so we need
// to give the top of the stack
char *stack_top = stack + STACK_SIZE;
char *stack_top = child_stack + STACK_SIZE;
// clone create thread
pid_t pid = clone(thread_start, stack_top, SIGCHLD, NULL);
pid_t pid = clone(thread_start, stack_top, CLONE_VM | SIGCHLD, NULL);
if (pid == -1) {
perror("clone");
exit(1);