mirror of
https://github.com/cs341-illinois/coursebook.git
synced 2026-10-02 08:04:38 +08:00
Merge pull request #241 from cs341-illinois/concerns-batch3
Fix remaining clear-cut future-concerns items (batch 3)
This commit is contained in:
+10
-8
@@ -355,7 +355,7 @@ In simplified terms, the descriptor needs to be closed, flushed, or read to its
|
||||
For a handle to become the active handle, the application shall ensure that the actions below are performed between the last use of the handle (the current active handle) and the first use of the second handle (the future active handle). The second handle then becomes the active handle. All activity by the application affecting the file offset on the first handle shall be suspended until it again becomes the active file handle. (If a stream function has as an underlying function one that affects the file offset, the stream function shall be considered to affect the file offset.)
|
||||
\end{quote}
|
||||
|
||||
Summarizing as if two file descriptors are actively being used, the behavior is undefined.
|
||||
In summary, if both processes actively use handles to the same open file description, the result is undefined.
|
||||
The other note is that after a fork, the library code must prepare the file descriptor as if the other process were to make the file active at any time.
|
||||
The last bullet point concerns itself with how a process prepares a file descriptor in our case.
|
||||
|
||||
@@ -589,7 +589,7 @@ While the API for most filesystems has stayed the same on POSIX over the years,
|
||||
|
||||
\subsection{Cutting Edge File systems}
|
||||
|
||||
There are a few filesystem hardware nowadays that are truly cutting edge.
|
||||
There are a few filesystem hardware technologies nowadays that are truly cutting edge.
|
||||
The one we'd briefly like to touch on is AMD's StoreMI.
|
||||
We aren't trying to sell AMD chipsets, but the featureset of StoreMI warrants a mention.
|
||||
|
||||
@@ -938,7 +938,8 @@ These examples are adapted from those.
|
||||
|
||||
\subsection{Sequentially Consistent}
|
||||
|
||||
Sequentially consistent is the simplest, least error-prone and most expensive model. This model says that any change that happens, all changes before it will be synchronized between all threads.
|
||||
Sequentially consistent is the simplest, least error-prone and most expensive model.
|
||||
This model says that there is a single total order of all operations, consistent with each thread's program order, that all threads agree on.
|
||||
|
||||
Suppose that \keyword{x} is atomic, \keyword{y} is an ordinary variable, and both start at 0.
|
||||
|
||||
@@ -982,7 +983,7 @@ This model was introduced so that there can be an Acquire/Release/Consume model
|
||||
There are a \textit{lot} of other methods of concurrency than described in this book.
|
||||
Posix threads are the finest grained thread construct, allowing for tight control of the threads and the CPU.
|
||||
Other languages have their abstractions.
|
||||
We'll talk about a language go that is similar to C in terms of simplicity and design, go or golang
|
||||
We'll talk about a language that is similar to C in terms of simplicity and design: Go (or golang).
|
||||
To get the 5 minute introduction, feel free to read \href{https://learnxinyminutes.com/docs/go/}{the learn x in y guide} for go.
|
||||
Here is how we create a "thread" in go.
|
||||
|
||||
@@ -1284,7 +1285,7 @@ The last bit of notation is that we will assume that the probability of getting
|
||||
\]
|
||||
|
||||
Which says that the scheduler needs to wait for all jobs with a higher priority and the same to go before a process can go.
|
||||
Imagine a series of FCFS queues that a process needs to wait your turn.
|
||||
Imagine a series of FCFS queues that a process needs to wait through before it gets its turn.
|
||||
Using Little's Law for different colored jobs and the formula above we can simplify this
|
||||
|
||||
\[
|
||||
@@ -1389,7 +1390,8 @@ We will also introduce an additional term $C_i$ which denotes the variation amon
|
||||
What happens?
|
||||
As always the proof is left to the reader.
|
||||
|
||||
\item Turnaround Time is the same formula $E[T] = E[S] + E[W]$. This means that given a distribution of jobs that has either low waiting time as described above, we will get low turnaround time -- we can't control the distribution of service times.
|
||||
\item Turnaround Time is the same formula $E[T] = E[S] + E[W]$.
|
||||
This means that given a distribution of jobs that has low waiting time as described above, we will get low turnaround time -- we can't control the distribution of service times.
|
||||
\end{enumerate}
|
||||
|
||||
\subsection{Preemptive Shortest Job First}
|
||||
@@ -1490,7 +1492,7 @@ The reasons are
|
||||
If the Internet Protocol receives a packet that is too big for the maximum size, it must chunk it up.
|
||||
TCP calculates how many datagrams that it needs to construct a packet and ensures that they are all transmitted and reconstructed at the end receiver.
|
||||
The reason that we barely use this feature is that if any fragment is lost, the entire packet is lost.
|
||||
Meaning that, assuming the probability of receiving a packet assuming each fragment is lost with an independent percentage, the probability of successfully sending a packet drops off exponentially as packet size increases.
|
||||
Assuming that each fragment is lost independently with the same probability, the probability of successfully sending a packet drops off exponentially as packet size increases.
|
||||
|
||||
As such, TCP slices its packets so that it fits inside one IP datagram.
|
||||
The only time that this applies is when sending UDP packets that are too big, but most people who are using UDP optimize and set the same packet size as well.
|
||||
@@ -1518,7 +1520,7 @@ So what are the benefits?
|
||||
In a high-performance server, this can easily happen 1000s of times a second.
|
||||
As such, having one system call to register and grab events saves the overhead of having a system call.
|
||||
\item The unified system call for all types.
|
||||
kqueue is the truest sense of underlying descriptor agnostic.
|
||||
kqueue is descriptor-agnostic in the truest sense.
|
||||
One can add files, sockets, pipes to it and get full or near full performance.
|
||||
You can add the same to epoll, but Linux's whole ecosystem with async file input-output has been messed up with \keyword{aio}, meaning that since there is no unified interface, you run into weird edge cases.
|
||||
\end{enumerate}
|
||||
|
||||
+24
-25
@@ -52,7 +52,7 @@ If it isn't, the processor requests a chunk of memory from the memory chip and s
|
||||
This is done because the l3 processor cache is roughly three times faster to reach than the memory in terms of time \cite[p. 22]{levinthal2009performance} though exact speeds will vary based on the clock speed and architecture.
|
||||
Naturally, this leads to problems because there are two different copies of the same value, in the cited paper this refers to an unshared line.
|
||||
This isn't a class about caching, but you should know how this could impact your code.
|
||||
A short but non-complete list could be
|
||||
A short but non-complete list could be:
|
||||
|
||||
\begin{enumerate}
|
||||
\item Race Conditions! If a value is stored in two different processor caches, then that value should be accessed by a single thread.
|
||||
@@ -104,7 +104,7 @@ Let's go through some of the common tools that you'll be working on and need to
|
||||
|
||||
\keyword{ssh} is short for the Secure Shell \cite{openbsd_ssh}.
|
||||
It is a network protocol that allows you to spawn a shell on a remote machine.
|
||||
Most of the time in this class you will need to ssh into your VM like this
|
||||
Most of the time in this class you will need to ssh into your VM like this:
|
||||
|
||||
\begin{lstlisting}[language=bash]
|
||||
$ ssh netid@sem-cs341-VM.cs.illinois.edu
|
||||
@@ -124,7 +124,7 @@ If you don't want to type your password out every time, you can generate an ssh
|
||||
If you still think that that is too much typing, you can always alias hosts.
|
||||
You may need to restart your VM or reload sshd for this to take effect.
|
||||
The config file is available on Linux and Mac distros.
|
||||
For Windows, you'll have to use the Windows Subsystem for Linux (WSL) or configure any aliases in PuTTY
|
||||
For Windows, you'll have to use the Windows Subsystem for Linux (WSL) or configure any aliases in PuTTY.
|
||||
|
||||
\begin{lstlisting}[language=bash]
|
||||
> cat ~/.ssh/config
|
||||
@@ -155,7 +155,7 @@ $ git push origin master
|
||||
Now to explain git well, you need to understand that git for our purposes will look like a linked list.
|
||||
You will always be at the head of master, and you will do the edit-add-commit-push loop. We have a separate branch on Github that we push feedback to under a specific branch which you can view on the Github website. The markdown file will have information on the test cases and results (like standard out).
|
||||
|
||||
Every so often git can break. Here is a list of commands you probably won't need to fix your repo
|
||||
Every so often git can break. Here is a list of commands you probably won't need to fix your repo:
|
||||
|
||||
\begin{enumerate}
|
||||
\item git-cherry-pick
|
||||
@@ -167,7 +167,7 @@ Every so often git can break. Here is a list of commands you probably won't need
|
||||
\item git-branch
|
||||
\end{enumerate}
|
||||
|
||||
If you are currently on a branch, and you don't see either
|
||||
Normally, \keyword{git status} shows that you are on a branch, with output like either
|
||||
|
||||
\begin{lstlisting}[language=bash]
|
||||
$ git status
|
||||
@@ -176,7 +176,7 @@ Your branch is up-to-date with 'origin/master'.
|
||||
nothing to commit, working directory clean
|
||||
\end{lstlisting}
|
||||
|
||||
or
|
||||
or:
|
||||
|
||||
\begin{lstlisting}[language=bash]
|
||||
$ git status
|
||||
@@ -192,7 +192,7 @@ Changes not staged for commit:
|
||||
no changes added to commit (use "git add" and/or "git commit -a")
|
||||
\end{lstlisting}
|
||||
|
||||
And something like
|
||||
If instead you see something like the following, with no branch at all:
|
||||
|
||||
\begin{lstlisting}[language=bash]
|
||||
$ git status
|
||||
@@ -201,12 +201,12 @@ nothing to commit, working directory clean
|
||||
\end{lstlisting}
|
||||
|
||||
Don't panic, but your repository may be in an unworkable state.
|
||||
If you aren't nearing a deadline, come to office hours or ask your question on Edstem, and we'd be happy to help.
|
||||
If you aren't nearing a deadline, come to office hours or ask your question on the class forum, and we'd be happy to help.
|
||||
In an emergency scenario, delete your repository and re-clone.
|
||||
\textbf{This will lose any local uncommitted changes. Make sure to copy any files you were working on outside the directory, remove and copy them back in.}
|
||||
|
||||
If you want to learn more about git, there are all but an endless number of tutorials and resources online that can help you.
|
||||
Here are some links that can help you out
|
||||
Here are some links that can help you out:
|
||||
|
||||
\begin{enumerate}
|
||||
\item \url{https://git-scm.com/docs/gittutorial}
|
||||
@@ -266,7 +266,7 @@ Many people will argue that editor gurus spend more time editing their editors t
|
||||
|
||||
Make your code modular using helper functions. If there is a repeated task (getting the pointers to contiguous blocks in the malloc MP, for example), make them helper functions.
|
||||
And make sure each function does one thing well so that you don't have to debug twice.
|
||||
Let's say that we are doing selection sort by finding the minimum element each iteration like so,
|
||||
Let's say that we are doing selection sort by finding the minimum element each iteration like so:
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
void selection_sort(int *a, long len){
|
||||
@@ -285,7 +285,7 @@ void selection_sort(int *a, long len){
|
||||
}
|
||||
\end{lstlisting}
|
||||
|
||||
Many can see the bug in the code, but it can help to refactor the above method into
|
||||
Many can see the bug in the code, but it can help to refactor the above method into these functions:
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
long max_index(int *a, long start, long end);
|
||||
@@ -473,7 +473,7 @@ $1 = 42
|
||||
\end{lstlisting}
|
||||
|
||||
You can also set breakpoints interactively from within gdb.
|
||||
Assume that we have no optimization and the line numbers are as follows
|
||||
Assume that we have no optimization and the line numbers are as follows:
|
||||
|
||||
\begin{lstlisting}[language=bash]
|
||||
1. int main() {
|
||||
@@ -500,7 +500,7 @@ $1 = 42
|
||||
|
||||
|
||||
We can also use gdb to check the content of different pieces of memory.
|
||||
For example,
|
||||
For example, consider this program:
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
int main() {
|
||||
@@ -509,7 +509,7 @@ int main() {
|
||||
}
|
||||
\end{lstlisting}
|
||||
|
||||
Compiled we get
|
||||
Compiling and running it, we get:
|
||||
|
||||
\begin{lstlisting}[language=bash]
|
||||
$ gcc main.c -g -o main && ./main
|
||||
@@ -517,7 +517,7 @@ $ Cat ZVQ� $
|
||||
\end{lstlisting}
|
||||
|
||||
|
||||
We can now use gdb to look at specific bytes of the string and reason about when the program should've stopped running
|
||||
We can now use gdb to look at specific bytes of the string and reason about when the program should've stopped running:
|
||||
|
||||
\begin{lstlisting}[language=bash]
|
||||
(gdb) l
|
||||
@@ -853,8 +853,7 @@ The virtual machine-in-your-browser and the videos you need for HW0 are here:
|
||||
|
||||
\url{http://cs-education.github.io/sys/}
|
||||
|
||||
Questions? Comments? Use the current semester's CS341 Edstem:
|
||||
\url{https://edstem.org/}
|
||||
Questions? Comments? Use the current semester's CS 341 Ed Discussion board, linked from the \href{https://cs341.cs.illinois.edu/}{course home page}.
|
||||
|
||||
The in-browser virtual machine runs entirely in JavaScript and is fastest in Chrome.
|
||||
Note the VM and any code you write is reset when you reload the page, \textbf{so copy your code to a separate document.}
|
||||
@@ -865,7 +864,7 @@ HW0 questions are below. Copy your answers into a text document because you'll n
|
||||
|
||||
\subsection{Chapter 1}
|
||||
|
||||
In which our intrepid hero battles standard out, standard error, file descriptors and writing to files
|
||||
In which our intrepid hero battles standard out, standard error, file descriptors and writing to files.
|
||||
|
||||
\begin{enumerate}
|
||||
\item \textbf{Hello, World! (system call style)} Write a program that uses \keyword{write()} to print out "Hi! My name is <Your Name>".
|
||||
@@ -884,7 +883,7 @@ The triangle should look like this, for n = 3:
|
||||
|
||||
\subsection{Chapter 2}
|
||||
|
||||
Sizing up C types and their limits, \keyword{int} and \keyword{char} arrays, and incrementing pointers
|
||||
Sizing up C types and their limits, \keyword{int} and \keyword{char} arrays, and incrementing pointers.
|
||||
|
||||
\begin{enumerate}
|
||||
\item How many bits are there in a byte?
|
||||
@@ -914,7 +913,7 @@ ssize_t str_len = strlen("Hello\0World");
|
||||
|
||||
\subsection{Chapter 3}
|
||||
|
||||
Program arguments, environment variables, and working with character arrays (strings)
|
||||
Program arguments, environment variables, and working with character arrays (strings).
|
||||
|
||||
\begin{enumerate}
|
||||
\item What are at least two ways to find the length of \keyword{argv}?
|
||||
@@ -931,7 +930,7 @@ What are the values of \keyword{sizeof(ptr)} and \keyword{sizeof(array)}? Why?
|
||||
|
||||
\subsection{Chapter 4}
|
||||
|
||||
Heap and stack memory, and working with structs
|
||||
Heap and stack memory, and working with structs.
|
||||
|
||||
\begin{enumerate}
|
||||
\item If I want to use data after the lifetime of the function it was created in ends, where should I put it? How do I put it there?
|
||||
@@ -978,7 +977,7 @@ Text input and output and parsing using \keyword{getchar}, \keyword{gets}, and \
|
||||
\subsection{C Development}
|
||||
|
||||
These are general tips for compiling and developing using a compiler and git.
|
||||
Some web searches will be useful here
|
||||
Some web searches will be useful here.
|
||||
|
||||
\begin{enumerate}
|
||||
\item What compiler flag is used to generate a debug build?
|
||||
@@ -1024,13 +1023,13 @@ Oh, and did I mention that this is an easy way to score points with your interns
|
||||
\item Am I following good programming practice? (i.e. encapsulation, functions to limit repetition, etc)
|
||||
\end{enumerate}
|
||||
|
||||
The biggest tip that we can give you when asking a question on the class forum if you want a swift answer is to \textbf{ask your question like you were trying to answer it}. Like before you ask a question, try to answer it yourself. If you are thinking about posting
|
||||
The biggest tip that we can give you when asking a question on the class forum if you want a swift answer is to \textbf{ask your question like you were trying to answer it}. Like before you ask a question, try to answer it yourself. If you are thinking about posting:
|
||||
|
||||
\begin{quote}
|
||||
Hi, My code got a 50% on the autograder. I tried testing it a lot, but couldn't get it to fail. Can you give me some hints as to the test cases
|
||||
Hi, My code got a 50\% on the autograder. I tried testing it a lot, but couldn't get it to fail. Can you give me some hints as to the test cases
|
||||
\end{quote}
|
||||
|
||||
Sounds good and courteous, but course staff would much much prefer a post resembling the following
|
||||
Sounds good and courteous, but course staff would much much prefer a post resembling the following:
|
||||
|
||||
\begin{quote}
|
||||
Hi, I recently failed test X, Y, Z which is about half the tests on this current assignment. I noticed that they all have something to do with networking and epoll, but couldn't figure out what was linking them together, or I may be completely off track. So to test my idea, I tried spawning 1000 clients with various get and put requests and verifying the files matched their originals. I couldn't get it to fail while running normally, the debug build, or valgrind or tsan. I have no warnings and none of the pre-syntax checks showed me anything. Could you tell me if my understanding of the failure is correct and what I could do to modify my tests to better reflect X, Y, Z? netid: bvenkat2
|
||||
|
||||
@@ -1139,7 +1139,7 @@ The one drawback is that you need more disks to have this setup, and there are m
|
||||
|
||||
Failure is common.
|
||||
Google reports 2-10\% of disks fail per year.
|
||||
Multiplying that by 60,000+ disks in a single warehouse.
|
||||
Multiplied across 60,000+ disks in a single warehouse, that is roughly 1,200 to 6,000 failed disks a year, or about 3 to 16 every day.
|
||||
Services must survive single disk, rack of servers, or whole data center failures.
|
||||
|
||||
\subsection{Solutions}
|
||||
@@ -1259,15 +1259,15 @@ Some questions to consider.
|
||||
|
||||
\begin{itemize}
|
||||
\item How would a program perform a write that goes across data block boundaries?
|
||||
\item How would a program perform a write after adding the offset would extend the length of the file?
|
||||
\item How would a program perform a write that starts inside the file but, once the offset is added, extends past the end of the file?
|
||||
\item How would a program perform a write where the offset is greater than the length of the original file?
|
||||
\end{itemize}
|
||||
|
||||
\subsubsection{Writing to directories}
|
||||
Performing a write to a directory implies that an inode needs to be added to a directory.
|
||||
If we pretend that the example above is a directory.
|
||||
Let's pretend that the example above is a directory.
|
||||
We know that we will be adding at most one directory entry at a time.
|
||||
Meaning that we have to have enough space for one directory entry in our data blocks.
|
||||
That means we need enough free space for one directory entry in our data blocks.
|
||||
Luckily the last data block that we have has enough free space.
|
||||
This means we need to find the number of the last data block as we did above, go to where the data ends, and write one directory entry.
|
||||
Don't forget to update the size of the directory so that the next creation doesn't overwrite your file!
|
||||
|
||||
@@ -16,9 +16,10 @@ file is the leftovers.
|
||||
- **These findings are unverified unless marked otherwise.** They were
|
||||
produced by an automated pass and will contain false positives. Confirm
|
||||
before acting, especially on technical claims.
|
||||
- Items that have since been fixed (PR #235 and the Tier 2 PR that
|
||||
followed it) have been removed, so line numbers may have drifted.
|
||||
Locate items by their quoted text.
|
||||
- Items that have since been fixed (PRs #234-#240 and the batch that
|
||||
followed them) have been removed, so line numbers may have drifted.
|
||||
Locate items by their quoted text. What remains needs an author's
|
||||
decision or new content.
|
||||
|
||||
---
|
||||
|
||||
@@ -31,110 +32,16 @@ but is currently discarded everywhere: the PDF is untagged, and pandoc 2.7
|
||||
(`\DocumentMetadata`) fails with the book's listings setup under TL2023.
|
||||
Pandoc 3.x does honour `alt=`, but then figures without it get empty alt,
|
||||
which the EPUB filters' `NoAltTagException` rejects, so the filters should
|
||||
fall back to the caption when pandoc is upgraded. Alt text has been added
|
||||
fall back to the caption when pandoc is upgraded (tracked in issue #238).
|
||||
Alt text has been added
|
||||
to all 48 figures,
|
||||
plus a sentence of prose wherever a figure carried facts the text did not;
|
||||
that prose is the only part that reaches readers of every format today.
|
||||
|
||||
---
|
||||
|
||||
## introduction
|
||||
|
||||
### introduction/introduction.tex:6 — dangling pronoun "It"
|
||||
|
||||
> "It is a message etched into our Alma Mater and makes up the DNA of our course staff."
|
||||
|
||||
The preceding sentence's subject is "we", not a message, so "It" has no clear antecedent — the intended referent is presumably the *belief* stated in line 5. Rewriting requires knowing the author's intent, so flagging rather than changing.
|
||||
|
||||
### introduction/introduction.tex:14 — possibly stale staff URL
|
||||
|
||||
> \href{http://cs341.cs.illinois.edu/staff}{CS 341 course staff}
|
||||
|
||||
Plain `http` (not `https`) and a course-site path that may have moved between semesters. A human should confirm the link still resolves.
|
||||
|
||||
### introduction/introduction.tex:17 — vague link text
|
||||
|
||||
> "This work is based on the original coursebook located \href{...}{at this url}."
|
||||
|
||||
I removed a duplicated "at" ("located at ... at this url"). The remaining link text "at this url" is still non-descriptive, which is an accessibility concern for screen readers; naming the target (e.g. the original SystemProgramming wiki) would be better, but that is a wording change, so left to a human.
|
||||
|
||||
### introduction/introduction.tex:20 — reference to "the duck"
|
||||
|
||||
> "Oh and the duck? Keep reading until synchronization :)."
|
||||
|
||||
Depends on a duck image/joke appearing in the synchronization chapter. Worth a human check that the referenced content still exists in the current build; there is no `\ref{}` tying the two together.
|
||||
|
||||
### introduction/introduction.tex:28 — AUTHORS.md included as a code listing
|
||||
|
||||
> \lstinputlisting[language=console]{AUTHORS.md}
|
||||
|
||||
The Authors section renders a markdown file in a monospaced listing environment. If AUTHORS.md is missing or moves, the build breaks silently in terms of content; also renders prose as code, which is an accessibility/presentation concern. Human decision.
|
||||
|
||||
---
|
||||
|
||||
## background
|
||||
|
||||
### Broken logic in the git-status troubleshooting flow — background/background.tex:168-201
|
||||
|
||||
"If you are currently on a branch, and you don't see either \<A\> or \<B\>" ... then line 193 continues "And something like \<C\>". The condition never resolves grammatically or logically: is the trigger *not* seeing A/B, or *seeing* C? As written a student can't tell what state means "don't panic, but your repository may be in an unworkable state". Needs an author rewrite.
|
||||
|
||||
### Generic Edstem link — background/background.tex:847-848
|
||||
|
||||
> "Use the current semester's CS341 Edstem: \url{https://edstem.org/}"
|
||||
|
||||
Left exactly as-is per instructions; noting only that it points at the site root rather than a course, which may be intentional.
|
||||
|
||||
### Inconsistent list-introduction punctuation — throughout
|
||||
|
||||
Several sentences that introduce an enumerate/lstlisting end without a period (e.g. lines 55, 126, 156, 207, 512, 972). This is consistent enough across the chapter to read as house style, so I left all of them alone rather than making a large punctuation-only diff.
|
||||
|
||||
---
|
||||
|
||||
## introc
|
||||
|
||||
### introc/language_facilities.tex:369 — ungrammatical struct definition
|
||||
"C-structs are contiguous regions of memory that one can access specific elements of each memory as if they were separate variables."
|
||||
The relative clause is broken; needs rewriting by someone who knows the intended sentence.
|
||||
|
||||
### introc/language_facilities.tex:497-498 — `void` / lvalue claim
|
||||
"The other use of \keyword{void} is when you are defining an \keyword{lvalue}." and "it can be promoted to any time to any other type."
|
||||
"any time" appears to be a typo for "any type", but the whole sentence (void* and lvalues) is technically confused, so I did not guess. Also "Pointer arithmetic with this pointer is undefined behavior" contradicts pointers.tex:147-148 which says gcc/clang permit it as a char*.
|
||||
|
||||
### introc/common_c_functions.tex:12 — broken sentence
|
||||
"know that most functions in C handle errors return oriented."
|
||||
Probably "handle errors in a return-oriented way". Needs an author's wording.
|
||||
|
||||
### introc/common_c_functions.tex:317 — broken sentence
|
||||
"The caller has to be careful from a valid 0 and an error."
|
||||
Presumably "has to distinguish a valid 0 from an error."
|
||||
|
||||
### introc/common_c_functions.tex:341 — dangling fragment
|
||||
"\keyword{memcpy} and \keyword{memmove} both in \keyword{string.h}?"
|
||||
This is not a sentence and the itemize ends on it. Possibly a leftover note ("Why are memcpy and memmove both in string.h?").
|
||||
|
||||
### introc/c_memory_model.tex:66-98 — figures have no alt text
|
||||
The three \includegraphics figures (memory_model_empty.eps, memory_model_length.eps, memory_model_full.eps) rely on captions only. The captions are descriptive, but there is no alt-text mechanism for screen readers.
|
||||
|
||||
### introc/pointers.tex:94 — confusing sentence
|
||||
"In addition to adding to an integer, pointers can be added to."
|
||||
Presumably "In addition to being able to add integers to integers, you can add an integer to a pointer." As written it is close to meaningless.
|
||||
|
||||
### introc/crash_course_introduction_to_c.tex:29 — flushing claim
|
||||
"If the newline isn't included, the buffer will not be flushed (i.e. the write will not complete immediately)." True only for a line-buffered stdout, and the buffer is still flushed at exit. common_c_functions.tex:99-101 states the nuanced version; this simplified claim may mislead.
|
||||
|
||||
### introc/crash_course_introduction_to_c.tex:112 — missing word
|
||||
"taking the sizeof the pointer and dividing it by the size of the first entry" — reads as if a word is missing ("the size of the pointer"). I left it because `sizeof` is being used as an operator name and a fix could change the technical reading.
|
||||
|
||||
---
|
||||
|
||||
## processes
|
||||
|
||||
### processes/processes.tex:142 — "starts at ... and starts at a constant size"
|
||||
"This section starts at the end of the text segment and starts at a constant size because the number of
|
||||
globals is known at compile time." The second "starts at" reads like it should be "stays at" / "has a
|
||||
constant size", but since this is a statement about segment layout I did not want to alter the meaning.
|
||||
(Compare line 163, which says the BSS "is also static in size".)
|
||||
|
||||
### processes/processes.tex:142 vs 127 — two different definitions of "program break"
|
||||
Line 127 says the program break is the top of the heap ("\keyword{malloc} may push the heap boundary --
|
||||
called the program break -- upward"); line 142 says "The end of the data segment is called the
|
||||
@@ -145,134 +52,27 @@ confuse students.
|
||||
"What is the difference between execs with a p and without a p? What does the operating system" — the
|
||||
second question has no verb, object, or terminal punctuation. I cannot guess the intended completion.
|
||||
|
||||
### processes/processes.tex:683-692 — person shifts between "your" and "its"
|
||||
"It is good practice to wait on your process' children. If a parent doesn't wait on your children they
|
||||
become ... If a long-running parent never waits for your children ... Having said that, a program doesn't
|
||||
always need to wait for your children! Your parent process can continue ..." The second-person "your"
|
||||
is attached to the parent process rather than the reader, which reads as an error, but fixing it means
|
||||
rewriting most of the paragraph, so I left it.
|
||||
|
||||
### processes/processes.tex:10 — dangling comparison
|
||||
"most systems that we'll be studying are almost POSIX compatible due more to political reasons." "due
|
||||
more to" invites a "than ..." that never arrives, and the claim itself (political reasons) is asserted
|
||||
with no context a student could use.
|
||||
|
||||
### processes/processes.tex:185,340,881 — figures have captions but no alt text
|
||||
`\includegraphics` of `address_space.eps`, `sleepsort_timing.eps`, and `fork_exec_wait.eps` carry only
|
||||
`\caption{}`. For an accessible PDF these need real alternative descriptions (the sleepsort timing
|
||||
diagram in particular carries information not present in its caption).
|
||||
|
||||
---
|
||||
|
||||
## malloc
|
||||
|
||||
### malloc/malloc.tex:9 — "use as its accord"
|
||||
"a contiguous series of addresses that the program can expand or contract and use as its accord". "as its accord" is not an English idiom; likely intended "as it sees fit" or "at its discretion". Needs an author decision on intended meaning rather than a guess.
|
||||
|
||||
### malloc/malloc.tex:83 — "these limitations" has no antecedent
|
||||
"An advanced discussion of these limitations is \href{...}{in this article}." The preceding sentence describes what `calloc` does; no limitations have been mentioned yet. A student cannot tell what limitations are meant. Also the linked host (locklessinc.com) may be dead — worth checking.
|
||||
|
||||
### malloc/malloc.tex:211 vs figure caption — "perfect-fit" vs "Best fit"
|
||||
Prose says "A perfect-fit strategy finds the smallest hole"; the figure caption immediately below says "Best fit finds an exact match", and the rest of the chapter (and the Topics list) uses "Best Fit". Terminology inconsistency that could confuse a student; renaming is an editorial call.
|
||||
|
||||
### malloc/malloc.tex:290 — Fibonacci heaps claim
|
||||
"Your heap could be represented with the max-heap data structure ... Using Fibonacci heaps, however, could be extremely inefficient." Fibonacci heaps have excellent amortized bounds; the claim as written is surprising and unexplained (presumably about constant factors / pointer overhead / cache behavior). Either justify or drop.
|
||||
|
||||
### malloc/malloc.tex:430-432 — broken quotation
|
||||
The `quote` block ends: "...a multiple of 16 on 64-bit systems." For example, if you need to calculate how many 16 byte units are required, don't forget to round up." There is a stray closing double-quote mid-block, and the "For example..." sentence is the book's own commentary sitting inside the glibc quotation. Also the quoted text is self-contradictory ("always a multiple of eight on most systems"). Fixing requires deciding where the quotation actually ends, and possibly re-checking the glibc manual wording.
|
||||
|
||||
### malloc/malloc.tex:487 — incomplete sentence
|
||||
"No more than 3 blocks will need to coalesce into a single block, and using a most recently used block scheme only one linked list entry." The second clause has no verb (presumably "...only one linked list entry needs to be updated"). Repairing it requires knowing the intended claim, so flagged rather than guessed.
|
||||
|
||||
### Figures — no alt text
|
||||
All figures (lines ~199-235, 306-319, 410-414, 475-479, 512-516, 530-534) use `\includegraphics` with a `\caption` only. The captions ("Malloc addition", "Free list good and bad coalesce") do not describe what the diagram shows, so a student using a screen reader or reading the text alone gets nothing. Accessibility improvement needs an author who knows the drawings.
|
||||
|
||||
---
|
||||
|
||||
## threads
|
||||
|
||||
### threads/threads.tex:205 — sentence fragment / duplicated "means"
|
||||
|
||||
```
|
||||
This means that the execution of the code is non-deterministic.
|
||||
Meaning that the same program can run multiple times and depending on how the kernel schedules the threads could produce inaccurate results.
|
||||
```
|
||||
|
||||
The second sentence is a fragment and repeats "means"; it also needs commas around the "depending on..." clause. Rewriting it is more than a mechanical fix.
|
||||
|
||||
### threads/threads.tex:230-231 — confusing register description
|
||||
|
||||
```
|
||||
We will assume that data is stored in the \keyword{eax} register.
|
||||
The code to increment is the following with no optimization (assume int\_ptr contains eax).
|
||||
```
|
||||
|
||||
"assume int_ptr contains eax" reverses the relationship, and the following assembly actually loads from `[rbp-4]`, not from a register holding `data`. Also the operation is a doubling, described as "increment". Needs an author's eye.
|
||||
|
||||
### threads/threads.tex:304 — description of the cast is inaccurate
|
||||
|
||||
```
|
||||
We will instead treat i as a pointer and cast it by value.
|
||||
```
|
||||
|
||||
The code passes the *value* of `i` cast to `void *`; "treat i as a pointer" is backwards, and "cast it by value" is not standard terminology. (The listing itself also uses `int data = ((int) ptr);`, which is implementation-defined on LP64 and normally warns.)
|
||||
|
||||
### threads/threads.tex:3 — epigraph
|
||||
|
||||
```
|
||||
\epigraph{If you think your programs were crashing before, wait until they crash ten times as fast}{}
|
||||
```
|
||||
|
||||
I inserted the missing verb ("programs crashing" -> "programs were crashing"). The original may have been intended as "your program's crashing"; flagging in case the author prefers that reading. No terminal punctuation, left as-is (epigraph style).
|
||||
|
||||
### threads/threads.tex:551 — awkward question
|
||||
|
||||
```
|
||||
What are a few things that threads share in a process? What are a few things that threads have different?
|
||||
```
|
||||
|
||||
"have different" is ungrammatical but the intended phrasing ("that differ between threads"?) is a judgement call, so left alone.
|
||||
|
||||
---
|
||||
|
||||
## synchronization
|
||||
|
||||
### Mutex description may be misleading
|
||||
|
||||
`synchronization/synchronization.tex:235-236` — "If a mutex is locked, the other threads will continue. It's only when a thread attempts to lock a mutex that is already locked, will the thread have to wait." The second sentence is ungrammatical (a mixed "It is only when… that…" / "Only when… will…" construction). Rewording touches a technical claim, so I left it.
|
||||
|
||||
### Confusing mutual-exclusion justification
|
||||
|
||||
`synchronization/synchronization.tex:404-405` — "How does this guarantee mutual exclusion? When working with atomics we are unsure! But in this simple example, we can because the thread that can successfully expect the lock to be UNLOCKED (0) and swap it…". The sentence has no clear main clause ("we can" what?) and "successfully expect" is odd. Technical passage, left alone.
|
||||
|
||||
### Semaphore-vs-mutex passage looks logically inverted
|
||||
|
||||
`synchronization/synchronization.tex:472` — "That is usually why a mutex is used to implement a semaphore and vice versa." reads as a non-sequitur after the warning about breaking the mutex abstraction. (The related line 527 claim about unlocking a mutex from another thread has been fixed.)
|
||||
|
||||
### Run-on sentence spanning a technical claim
|
||||
|
||||
`synchronization/synchronization.tex:531` — "\keyword{sem\_post} is one of a handful of functions that can be correctly used inside a signal handler \keyword{pthread\_mutex\_unlock} is not." Two sentences fused with no punctuation. I did not insert punctuation because the fix (semicolon vs. period vs. "whereas") changes emphasis on an async-signal-safety claim; a one-character insert is easy for the author.
|
||||
|
||||
### Structural: "Sketch #1" is never analysed; text jumps to Sketch #2
|
||||
|
||||
`synchronization/synchronization.tex:855-877` — the listing is labelled `// Sketch #1` and is syntactically broken (a `push` nested inside `pop`, unbalanced braces), and the very next paragraph starts "Sketch \#2 has implemented the \keyword{post} too early." Sketch #1 is never discussed. Reads like a missing paragraph.
|
||||
|
||||
### Bounded-wait definition is awkward
|
||||
|
||||
`synchronization/synchronization.tex:1034` — "A thread/process cannot get superseded by another thread infinite amounts of time." Probably "an infinite number of times". Left because the fix is a judgement call on the intended definition.
|
||||
|
||||
### Garbled question
|
||||
|
||||
`synchronization/synchronization.tex:2267` — "How might the above be a producer consumer problem be used in the above section?" Doubled "be" and duplicated "the above"; the intended question is unclear.
|
||||
|
||||
### Possibly incomplete question prompt
|
||||
|
||||
`synchronization/synchronization.tex:2357` — "Remember in addition to mutual exclusion, a mutex can only ever be unlocked by the thread who called it." "the thread who called it" is missing what was called (presumably "the thread that locked it").
|
||||
|
||||
### Figure without alt text
|
||||
|
||||
`synchronization/synchronization.tex:1792-1795` — `\includegraphics{synchronization/drawings/ring_buffer.eps}` with only `\caption{Ring Buffer Visualization}`. No descriptive alternative text for a figure carrying real content (index wrap-around). Accessibility.
|
||||
|
||||
---
|
||||
|
||||
## deadlock
|
||||
@@ -307,9 +107,6 @@ Ungrammatical inside a proof; the intended sense is probably "acting under the p
|
||||
### deadlock/deadlock.tex:378-383 — Dijkstra proof reduction
|
||||
Line 378 "If the last philosopher $p_{n-1}$ holds the first lock meaning the previous philosopher $p_{n-2}$ is waiting on $r_{n-1}$ meaning $r_{n-2}$ is available" is a run-on with no main verb, and line 380 concludes "we now have $n$ resources but only $n-1$ philosophers" without stating which philosopher was removed. Also line 379 uses "her" while the rest of the passage uses "he/his". The substance of the reduction needs an author's check.
|
||||
|
||||
### deadlock/deadlock.tex:24-29, 68-72, 191-195, 227-231, 268-272, 303-307, 345-349, 387-391 — figures have no alt text
|
||||
Eight `\includegraphics` calls, none with alt text; several carry load-bearing content (the deadlock cycle, the livelock time evolution, the arbitrator diagram). Only three of the eight are referenced from the prose at all (`ragfigure` is the sole `\label`), so a reader relying on a screen reader loses the content entirely. Needs an accessibility decision at the book level.
|
||||
|
||||
---
|
||||
|
||||
## ipc
|
||||
@@ -319,41 +116,13 @@ Eight `\includegraphics` calls, none with alt text; several carry load-bearing c
|
||||
|
||||
Ungrammatical and technically muddled: it says "two intermediate levels" when the surrounding text (line 186-188) describes two *sub-tables*, and "4KiB for the two intermediate levels" reads as 4KiB total while 2+4+4=10KiB implies 4KiB each. A human should restate this sentence.
|
||||
|
||||
### ipc/ipc.tex:217 — "read and write" in MMU pseudocode
|
||||
> "get the physical frame from the TLB and perform the read and write."
|
||||
|
||||
Should presumably be "the read or write" (a single access is either). Left alone as it is inside the algorithm description.
|
||||
|
||||
### ipc/ipc.tex:223 — Broken pseudocode step
|
||||
> "If so then do the dereference provide the address, cache the results in the TLB"
|
||||
|
||||
Run-on with words apparently missing ("provide the address" is unattached) and no terminal punctuation. Intent unclear, needs an author.
|
||||
|
||||
### ipc/ipc.tex:248 — Sentence ends with a dangling verb
|
||||
> "it all depends on if your hardware says that a program can access."
|
||||
|
||||
"can access" has no object (access what — that page?). Needs an author to complete.
|
||||
|
||||
### ipc/ipc.tex:645 — "your special byte" mixes person
|
||||
> "a program could write your special byte (e.g.~0xff)"
|
||||
|
||||
Probably "a special byte". Left as-is since the chapter deliberately mixes second person elsewhere.
|
||||
|
||||
### ipc/ipc.tex:955,980 — "this"/"That quirk" with no antecedent
|
||||
Section "Determining File Length" opens "using fseek and ftell is a simple way to accomplish this" (no prior referent), and section "Use stat instead" opens "This only works on some architectures and compilers. That quirk is that longs only need to be 4 Bytes big" — the quirk is named only after it is referred to. Reads as if an introductory sentence was lost.
|
||||
|
||||
### ipc/ipc.tex — Figures have no alt text
|
||||
All figures (lines 72-76, 95-99, 104-108, 112-116, 143-147, 161-165, 169-173, 506-510) carry only short `\caption{}` text such as "Splitting Address" and "One level dereference". For an accessible PDF these diagrams — which carry the core address-translation explanation — need real descriptions.
|
||||
|
||||
---
|
||||
|
||||
## scheduling
|
||||
|
||||
### "Unless otherwise stated" is a dangling fragment
|
||||
`scheduling/scheduling.tex:134`
|
||||
|
||||
The line introducing the shared example process list is just `Unless otherwise stated` with no verb and no terminal punctuation. Presumably intended as something like "Unless otherwise stated, the following processes are used in each example:". I did not guess at the intended wording.
|
||||
|
||||
### Cross-reference to "the appendix and the section conceptually scheduling" is informal/unverifiable
|
||||
`scheduling/scheduling.tex:331`
|
||||
|
||||
@@ -361,77 +130,14 @@ The line introducing the shared example process list is just `Unless otherwise s
|
||||
|
||||
There is no `\ref`/`\label` here, and "the section conceptually scheduling" does not read like an actual section title. A human should confirm the target exists and ideally replace this with a real `\ref{}`. The sentence also has no terminal period, but I left it since the whole line may be rewritten.
|
||||
|
||||
### Stray capitalization "Convoy Behind them"
|
||||
`scheduling/scheduling.tex:120`
|
||||
|
||||
> ...leaving all other processes with potentially smaller resource needs following like a Convoy Behind them.
|
||||
|
||||
"Behind" is capitalized mid-sentence for no apparent reason. It may be deliberate emphasis in this book's informal voice, so I left it. Also note "Convoy effect" (line 260) vs "Convoy Effect" (line 57) vs "convoy effect" (lines 118, 120, 262, 354) are inconsistently capitalized throughout.
|
||||
|
||||
### Figures have captions but no alt text
|
||||
`scheduling/scheduling.tex:146-150, 193-197, 237-241, 287-291`
|
||||
|
||||
All four `\includegraphics` calls (sjf.eps, psjf.eps, fcfs.eps, rr.eps) carry only short captions such as "Shortest job first scheduling". The Gantt-chart content — arrival times, ordering, and the resulting timeline — exists only in the image, so a student using a screen reader gets none of it. Worth adding descriptive alt text or an in-text summary of each chart.
|
||||
|
||||
### "with a high priority" where "higher" is likely meant
|
||||
`scheduling/scheduling.tex:66`
|
||||
|
||||
> Thus once a process is scheduled it will continue even if another process with a high priority appears on the ready queue.
|
||||
|
||||
The point being made is about a process of *higher* priority than the running one. Reads as a wording slip rather than a plain grammar error, so I left it for a human.
|
||||
|
||||
---
|
||||
|
||||
## networking
|
||||
|
||||
### Garbled IPv4 address-splitting sentence — networking/networking.tex:70
|
||||
|
||||
"Conceptually the source and destination addresses can be split into two: a network number the upper bits and lower bits represent a particular host number on that network." The sentence has no working structure and a student cannot extract the network/host split from it. Rewriting requires deciding what was meant, so I left it.
|
||||
|
||||
### Confusing IPv6 address-notation description — networking/networking.tex:74-75
|
||||
|
||||
"We write IPv6 addresses in a sequence of eight, four hexadecimal delimiters like \"1F45:0000:...\"". "eight, four hexadecimal delimiters" is not meaningful — presumably "eight groups of four hexadecimal digits". Also "Since that can get unruly, we can omit the zeros \"1F45::\"" understates the `::`-may-appear-once rule. Technical wording, so left for a human.
|
||||
|
||||
### "Ports" bullet says socket where it means port — networking/networking.tex:252-253
|
||||
|
||||
"TCP gives the programmer a set of virtual sockets. Clients specify the socket that you want the packet sent to". The concept being introduced is the *port*; calling it a socket here conflicts with the socket API introduced later and will confuse students.
|
||||
|
||||
### "High performance and error-prone code won't even assume that!" — networking/networking.tex:244
|
||||
|
||||
Unclear as written — presumably means high-performance / error-tolerant code should not assume delivery. As phrased ("error-prone code") it reads as praising buggy code. Needs an author decision.
|
||||
|
||||
### Garbled HTTP body description — networking/networking.tex:557
|
||||
|
||||
"The actual body of the request delimited by two new lines. The body of the request is either if the size is specified or until the receiver closes their connection." The second sentence is missing its predicate ("either read until the specified length..."). Needs an author rewrite.
|
||||
|
||||
### HTTP version/RFC currency — networking/networking.tex:574
|
||||
|
||||
"RFC 7231 has the most current specifications on the most common HTTP method today". RFC 7231 was obsoleted by RFC 9110 (HTTP Semantics, 2022), and the chapter's examples are all HTTP/1.0 while HTTP/1.1 and HTTP/2/3 dominate. A human should decide how much to update. Also line 553, "the HTTP/1.0 method" should probably be "protocol"/"version".
|
||||
|
||||
### "There are a variety of function calls available to send UDP sockets" — networking/networking.tex:902
|
||||
|
||||
You send *packets*, not sockets. Likely "to send data over UDP sockets". I could not fix it without guessing the intent.
|
||||
|
||||
### Garbled UDP-vs-TCP efficiency sentence — networking/networking.tex:825
|
||||
|
||||
"TCP has \textit{decades} of optimization, meaning your protocol for its use cases needs to be more efficient that to be more beneficial to use it." Not parseable; needs an author rewrite (also contains a then/than-adjacent "that").
|
||||
|
||||
### Server stub sentence missing a word — networking/networking.tex:1361
|
||||
|
||||
"unmarshal the request into a valid in-memory data call the underlying implementation and send the result back". Probably "into a valid in-memory representation, call the underlying implementation, and send...". Comma/word insertion needs the author's intent.
|
||||
|
||||
### Interface-Description-Language sentence loses its subject — networking/networking.tex:1372
|
||||
|
||||
"Writing stub code by hand is painful, tedious, error-prone, difficult to maintain and difficult to reverse engineer the wire protocol from the implemented code." The final clause does not attach to the list.
|
||||
|
||||
### Comma splice left as-is — networking/networking.tex:1319
|
||||
|
||||
"To marshal a linked list, it is unnecessary to send the link pointers, stream the values." Reads as a splice; the fix ("instead, stream the values") is a wording choice so I left it.
|
||||
|
||||
### Figures have no alt text — networking/networking.tex:88-92, 236-240
|
||||
|
||||
Both `\includegraphics` figures (`ipv6_datagram.eps`, `tcp_header.eps`) carry only captions ("IPv6 Datagram divisibility", "Extra: TCP Header Specification") and no textual description. The IPv6 caption in particular does not explain what the diagram shows. Accessibility issue for screen-reader users.
|
||||
|
||||
---
|
||||
|
||||
## filesystems
|
||||
@@ -442,12 +148,6 @@ Both `\includegraphics` figures (`ipv6_datagram.eps`, `tcp_header.eps`) carry on
|
||||
|
||||
Unclear what the baseline is (three times slower than a direct block? than a single indirect block?), and the factor is asserted without justification. Ambiguous enough to need an author.
|
||||
|
||||
### Sentence fragment in the Google disk-failure statistics — filesystems.tex:1142
|
||||
|
||||
> "Multiplying that by 60,000+ disks in a single warehouse."
|
||||
|
||||
No main verb, and the conclusion (how many failures per day that implies) is never stated — a student cannot finish the arithmetic from what is given. Also worth a date check on the "2-10\% of disks fail per year" figure.
|
||||
|
||||
### RAID-10 description is hard to follow — filesystems.tex:1101-1105
|
||||
|
||||
> "This means you would get roughly the same speed from the slowdowns but now any one disk can fail and you can recover that disk."
|
||||
@@ -466,68 +166,6 @@ No main verb, and the conclusion (how many failures per day that implies) is nev
|
||||
|
||||
This should almost certainly be the 2nd *data block* (the inode is a single object here), and the next sentence then says "go to the $5$th data block", which does not obviously follow from "2nd". Since the whole passage depends on the figure at filesystems.tex:1176, an author who can see the figure should reconcile the numbering.
|
||||
|
||||
### Follow-up question is garbled — filesystems.tex:1262
|
||||
|
||||
> "How would a program perform a write after adding the offset would extend the length of the file?"
|
||||
|
||||
Not a grammatical sentence, and it is unclear how it differs from the next question ("offset is greater than the length of the original file"). Left alone because the intended meaning is genuinely ambiguous.
|
||||
|
||||
### Figure has no alt text — filesystems.tex:1174-1178
|
||||
|
||||
`\includegraphics{filesystems/images/sample_file.png}` with caption "Sample file filling up". The whole "Simple Filesystem Model" section (file size bounds, reads, writes) is written entirely against this image — a student using a screen reader, or reading the text alone, cannot follow any of the worked calculations. Adding a textual description of the inode's block pointers would fix this.
|
||||
|
||||
### Fragment in "Writing to directories" — filesystems.tex:1268-1270
|
||||
|
||||
> "If we pretend that the example above is a directory. We know that we will be adding at most one directory entry at a time. Meaning that we have to have enough space for one directory entry in our data blocks."
|
||||
|
||||
Two sentence fragments in a row ("If we pretend..." with no main clause, and "Meaning that..."). Fixing them requires deciding what the sentences were meant to join to, so left to an author.
|
||||
|
||||
---
|
||||
|
||||
## signals
|
||||
|
||||
### signals.tex:7 — "Sometimes, a program can choose to ignore events which is supported."
|
||||
Circular/confusing as written; it is unclear whether the point is that ignoring is a supported disposition, or that only some signals may be ignored (SIGKILL/SIGSTOP cannot). A student would benefit from the caveat being stated explicitly here.
|
||||
|
||||
### signals.tex:58-62 — figure has no alt text
|
||||
`\includegraphics{signals/drawings/signal_lifecycle.eps}` with caption "Signal lifecycle diagram" only. The caption does not convey the lifecycle content to a reader using a screen reader, and the surrounding text (line 56, "As a flowchart") does not describe it either.
|
||||
|
||||
### signals.tex:262 — sig_atomic_t range claim
|
||||
"can be as small as a \keyword{char} and only able to represent (-127 to 127) values" — a technical claim about limits (C requires at least SIG_ATOMIC_MIN/MAX coverage) that I did not want to touch.
|
||||
|
||||
### signals.tex:46 — "the process' signal mask"
|
||||
Possessive of a singular noun ending in s-sound written as `process'`; elsewhere I normalized "processes mask" to "process's mask" (lines 431-432). Left line 46 alone to avoid churn, but the book should pick one convention.
|
||||
|
||||
---
|
||||
|
||||
## security
|
||||
|
||||
### Step 2 refers to an antecedent that doesn't exist — security/security.tex:37
|
||||
|
||||
> "First, you should determine if your use is intended or unintended or somewhere in the middle -- get a decision from them."
|
||||
|
||||
"them" has no antecedent in the sentence (presumably the system's owners/developers). I fixed "for them" -> "from them" but the referent is still dangling.
|
||||
|
||||
### "In lieu" used without an object — security/security.tex:51
|
||||
|
||||
> "In lieu, you must be able to say that you reacted as a ``reasonable'' engineer would react."
|
||||
|
||||
"In lieu" requires "of X"; the intended phrase is probably "In lieu of that" or "Instead". Left alone as it may be deliberate shorthand.
|
||||
|
||||
### "each user has a certain set of permissions that they can do" — security/security.tex:227
|
||||
|
||||
Grammatically mismatched ("permissions ... do") and conflates capabilities with permissions. Suggest "a certain set of capabilities" / "set of actions they are permitted to perform", but the wording sits inside a technical definition, so leaving to a human.
|
||||
|
||||
### DNS trust sentence is confusing — security/security.tex:324
|
||||
|
||||
> "One just has to trust the DNS server gave a reasonable response which is almost always the incorrect answer."
|
||||
|
||||
Unclear what "the incorrect answer" refers to — the DNS response, or the decision to trust it. Reads as a garbled sentence; needs the author's intent.
|
||||
|
||||
### Review question 381 vs body text — security/security.tex:336 and 381
|
||||
|
||||
Line 336 already states "Distributed Denial of Service is the hardest form of attack to stop", which answers review question 10 ("Which is harder to defend against: Syn-Flooding or Distributed Denial of Service?") outright. Intentional? Possibly fine, but flagging.
|
||||
|
||||
---
|
||||
|
||||
## review
|
||||
@@ -535,79 +173,27 @@ Line 336 already states "Distributed Denial of Service is the hardest form of at
|
||||
### review/review.tex:135 — truncated bonus question
|
||||
The item ends: "Bonus: How would you make this code more robust or able to cope with?" The sentence is cut off ("cope with" what — a long `mesg`? a `malloc` failure?). The same sentence also says "val as a double val", which looks like a duplicated word but might be intentional shorthand. Both need an author who knows the intended question; rewriting could change what is being asked.
|
||||
|
||||
### review/review.tex:197-203 — question text is split across a code listing
|
||||
"When would a trivial malloc implementation" / listing / "be acceptable?" I lowercased the stray capital "Be", but the sentence still reads oddly when the listing is set as a display block. A human may prefer to reword (e.g. "When would the trivial malloc implementation shown below be acceptable?").
|
||||
|
||||
### review/review.tex:322-337 — two items describe one problem, and item 2 has no question
|
||||
Item at 322 sets up the graph/`shortest`/`set_edge` scenario and ends at line 335 with a requirement statement but no explicit question ("For performance, multiple threads must be able to call \keyword{shortest} at the same time..."). The next `\item` (337) then asks for the reader-writer implementation of the same scenario. These probably should be one item, or item 1 needs an actual question sentence.
|
||||
|
||||
### review/review.tex:519 — chmod question sentence is ungrammatical
|
||||
"...so that the owner can read, write, and execute permissions the group can read and everyone else has no access." The verb "can" does not fit "permissions", and there is no punctuation separating the owner clause from the group clause. I only fixed the missing spaces after the commas; the rest is a rewrite that touches what the question asks (intended answer is presumably `chmod 740`), so a human should word it.
|
||||
|
||||
### review/review.tex:513,515,585 — space before question mark
|
||||
Three items have "notes.txt} ?" / "listen accept ?" with a space before the "?". Left alone as it may be a deliberate consequence of the `\keyword{}` macro spacing, but a human may want them tightened.
|
||||
|
||||
---
|
||||
|
||||
## honors
|
||||
|
||||
### honors/kernel.tex:11-13 — Windows/Darwin sentence was structurally broken
|
||||
Original text read as two fragments: "...the Windows kernel, which we won't talk about too much in this chapter." followed by a new line beginning "or \keyword{Darwin}, the UNIX-like kernel for macOS...". I joined them with a comma (minimal fix), but the result now reads as "we won't talk about Windows or Darwin", which may not be the intended meaning — the original may have lost a clause such as "Others may have used XNU or Darwin". Please confirm the intended sentence.
|
||||
|
||||
---
|
||||
|
||||
## appendix
|
||||
|
||||
### appendix/appendix.tex:358 — garbled sentence in the Fork-FILE explanation
|
||||
|
||||
> "Summarizing as if two file descriptors are actively being used, the behavior is undefined."
|
||||
|
||||
"Summarizing as" is not grammatical, and it is unclear whether the intended meaning is "Summarizing: if two file descriptors are actively being used..." or something narrower (POSIX's condition is about *handles* to the same open file description, not any two descriptors). Because the precise POSIX claim matters, a human should decide the wording.
|
||||
|
||||
### appendix/appendix.tex:442 — truncated sentence
|
||||
|
||||
> "Also, the messages now encounter additional overhead for serializing and deserializing or at the least."
|
||||
|
||||
The sentence ends mid-thought ("or at the least" what?). Needs the author to supply the missing clause.
|
||||
|
||||
### appendix/appendix.tex:592 — "a few filesystem hardware"
|
||||
|
||||
> "There are a few filesystem hardware nowadays that are truly cutting edge."
|
||||
|
||||
Ungrammatical count/mass mismatch; likely intended "a few filesystem hardware technologies" or "a few filesystems". Word choice is the author's call, and the section then discusses StoreMI (a hardware/software caching product), so the right noun is ambiguous.
|
||||
|
||||
### appendix/appendix.tex:639 — stray one-word paragraph "Yes"
|
||||
|
||||
Section `\subsection{Implementing Software Mutex}` opens with a bare line reading `Yes` before "With a bit of searching, it is possible to find it in production...". This looks like the answer to a question that was deleted (probably "Is Peterson's algorithm ever used in practice?"). The paragraph also has an unexplained "it". A human should restore the missing question or delete the line.
|
||||
|
||||
### appendix/appendix.tex:943 — Sequential consistency definition is hard to parse
|
||||
|
||||
> "This model says that any change that happens, all changes before it will be synchronized between all threads."
|
||||
|
||||
Grammatically broken and technically imprecise (sequential consistency is about a single total order of operations consistent with program order). Needs an author rewrite.
|
||||
|
||||
### appendix/appendix.tex:986 — duplicated/garbled clause introducing Go
|
||||
|
||||
> "We'll talk about a language go that is similar to C in terms of simplicity and design, go or golang"
|
||||
|
||||
The language name appears three times and the sentence has no terminal punctuation. Probably intended: "We'll talk about a language similar to C in terms of simplicity and design: Go (or golang)." Left alone because it is a rewrite, not a typo fix.
|
||||
|
||||
### appendix/appendix.tex:1126 vs 1317 — inconsistent symbol for maximum run time
|
||||
|
||||
Line 1126 says "Let the maximum amount of time that a process runs be equal to $S$", but line 1317 says "$T$ is the maximum amount of time a process can run for". Meanwhile $S$ is used throughout as the *service time* random variable. A student following the derivations would be confused; needs the author to pick one symbol.
|
||||
|
||||
### appendix/appendix.tex:1288 — broken sentence
|
||||
|
||||
> "Imagine a series of FCFS queues that a process needs to wait your turn."
|
||||
|
||||
Mixes third and second person and is missing words. Rewrite needed.
|
||||
|
||||
### appendix/appendix.tex:1393 — dangling "either"
|
||||
|
||||
> "given a distribution of jobs that has either low waiting time as described above"
|
||||
|
||||
"Either" has no second alternative. Likely a dropped clause.
|
||||
|
||||
### appendix/appendix.tex:1461 — orphan fragment in the routing list
|
||||
|
||||
> "These protocols are meant to be fast and more trusting because all computers, switches, and routers are part of an ISP.
|
||||
@@ -615,43 +201,5 @@ Mixes third and second person and is missing words. Rewrite needed.
|
||||
|
||||
The last line is a lowercase sentence fragment with no context — apparently a leftover from an edit. Needs the author to restore or delete.
|
||||
|
||||
### appendix/appendix.tex:1494 — duplicated "assuming" and unclear claim
|
||||
|
||||
> "assuming the probability of receiving a packet assuming each fragment is lost with an independent percentage, the probability of successfully sending a packet drops off exponentially as packet size increases"
|
||||
|
||||
The sentence is garbled; I could not tell which "assuming" to drop or how the independence assumption was meant to be phrased.
|
||||
|
||||
### appendix/appendix.tex:1522 — garbled sentence about kqueue
|
||||
|
||||
> "kqueue is the truest sense of underlying descriptor agnostic."
|
||||
|
||||
Not a grammatical sentence; probably intended "kqueue is descriptor-agnostic in the truest sense." Rewrite needed.
|
||||
|
||||
### appendix/appendix.tex:171-198, 1409-1413 — figures have no alt text
|
||||
|
||||
`\includegraphics` for `struct_clean.eps`, `struct_slop.eps`, and `ip_datagram.eps` have captions ("Six box struct", "IP Datagram divisibility") but no descriptive alternative text. The IP datagram figure in particular carries information not otherwise in the text. Accessibility decision for the author.
|
||||
|
||||
---
|
||||
|
||||
## post_mortems
|
||||
|
||||
### AT&T 1990: "operable when they weren't" may be backwards
|
||||
`post_mortems/post_mortems.tex:215` — "A series of network delays that caused some telephone switches across the country to think that other switches were operable when they weren't." The usual account of the January 1990 AT&T collapse is that a switch went down for maintenance, and the *recovery* message it sent when coming back up crashed its neighbours via a misplaced `break` in C code — i.e. switches wrongly concluded peers were *inoperable*/failing. Either wording direction is a factual claim about a real incident, so I left it. Also note this sentence is a fragment ("A series of network delays that caused...") with no main verb.
|
||||
|
||||
### Appnexus double-free description is hard to follow
|
||||
`post_mortems/post_mortems.tex:203` — "This is fine until two threads try to delete the same object at once, adding to the list twice. After less time, one of the objects was deleted, the delete was announced to other computers." "After less time" is meaningless as written, and the causal chain from double-add to outage is not explained. A student cannot reconstruct the bug from this. Needs a rewrite by someone who knows the incident.
|
||||
|
||||
### Mars Pathfinder paragraph: run-on and tense-mixing
|
||||
`post_mortems/post_mortems.tex:74,76` — Line 74 mixes past and present ("The finder uses a single bus...", "if an interrupt happened ... and a task is running and a task is to be scheduled"). Line 76 is a comma splice: "The pattern that caused everything to start failing was the data collection thread starts writing to the bus, the information bus thread is waiting on the data." Fixing this properly means restructuring sentences, which is beyond a low-risk copy-edit. Note the classic name for this bug — priority inversion — is never stated, which is the one term a student would want.
|
||||
|
||||
### Sentence fragment in the Sony rootkit section
|
||||
`post_mortems/post_mortems.tex:156` — "What websites visited, what clicks or keys typed etc." has no verb. It reads as a deliberate telegraphic aside in an informal chapter, so I left it, but a human may want "What websites are visited, what clicks or keys are typed, etc."
|
||||
|
||||
### Meltdown and Spectre sections are stubs with unverifiable pointers
|
||||
`post_mortems/post_mortems.tex:60-66` — "There is an example of this in the background section." / "Check in the security section." These are prose pointers, not `\ref{}`s, so they cannot be checked by the build and will silently rot if chapters are renamed or reordered. Consider real `\ref{}`s, or content.
|
||||
|
||||
### No `\label{}` anywhere in the chapter
|
||||
`post_mortems/post_mortems.tex` (whole file) — The chapter and its ~16 sections define no labels, so nothing elsewhere in the book can cross-reference an individual post-mortem. Structural, and a human's call.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -9,7 +9,7 @@ In the shell, \keyword{man -S2 open} or \keyword{man -S3 printf}.
|
||||
|
||||
\subsection{Handling Errors}
|
||||
|
||||
Before we get into the nitty gritty of all the functions, know that most functions in C handle errors return oriented.
|
||||
Before we get into the nitty gritty of all the functions, know that most functions in C report errors through their return value.
|
||||
This is at odds with programming languages like C++ or Java where the errors are handled with exceptions.
|
||||
There are a number of arguments against exceptions.
|
||||
|
||||
@@ -327,7 +327,7 @@ tricky
|
||||
Be careful though!
|
||||
Error handling is tricky because the function won't return an error code.
|
||||
If passed an invalid number string, it will return 0.
|
||||
The caller has to be careful from a valid 0 and an error.
|
||||
The caller has to be careful to distinguish a valid 0 from an error.
|
||||
This often involves an errno trampoline as shown below.
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
@@ -351,7 +351,7 @@ tricky
|
||||
Consider the safer version \keyword{memmove}.
|
||||
|
||||
\item \keyword{void *memmove(void *dest, const void *src, size\_t n)} does the same thing as above, but if the memory regions overlap then it is guaranteed that all the bytes will get copied over correctly.
|
||||
\keyword{memcpy} and \keyword{memmove} both in \keyword{string.h}?
|
||||
Both \keyword{memcpy} and \keyword{memmove} are declared in \keyword{string.h}.
|
||||
\end{itemize}
|
||||
|
||||
|
||||
|
||||
@@ -27,7 +27,8 @@ int main(void) {
|
||||
\keyword{printf} is defined as a part of \keyword{stdio.h}.
|
||||
The function has been compiled and lives somewhere else on our machine - the location of the C standard library.
|
||||
Just remember to include the header and call the function with the appropriate parameters (a string literal \keyword{"Hello World\textbackslash n"}).
|
||||
If the newline isn't included, the buffer will not be flushed (i.e. the write will not complete immediately).
|
||||
When standard output is a terminal it is usually line buffered, so without the newline the text may sit in the buffer and not appear immediately.
|
||||
The buffer is still flushed when the program exits normally.
|
||||
\item \keyword{return 0}.
|
||||
\keyword{main} has to return an integer.
|
||||
By convention, \keyword{return 0} means success and anything else means failure.
|
||||
@@ -111,4 +112,4 @@ int* dynamic_array = malloc(10); // ARRAY_LENGTH(dynamic_array) = 2 or 1 consist
|
||||
|
||||
What is wrong with the macro?
|
||||
Well, it works if a static array is passed in because \keyword{sizeof} a static array returns the number of bytes that array takes up and dividing it by the \keyword{sizeof(an\_element)} would give the number of entries.
|
||||
But if passed a pointer to a piece of memory, taking the sizeof the pointer and dividing it by the size of the first entry won't always give us the size of the array.
|
||||
But if passed a pointer to a piece of memory, taking the \keyword{sizeof} of the pointer and dividing it by the size of the first entry won't always give us the size of the array.
|
||||
|
||||
@@ -366,7 +366,7 @@ char *print_time(void) {
|
||||
\end{lstlisting}
|
||||
|
||||
\item \keyword{struct} is a keyword that allows you to pair multiple types together into a new structure.
|
||||
C-structs are contiguous regions of memory that one can access specific elements of each memory as if they were separate variables.
|
||||
C structs are contiguous regions of memory whose individual members you can access as if they were separate variables.
|
||||
Note that there might be padding between elements, such that each variable is memory-aligned (starts at a memory address that is a multiple of its size).
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
@@ -494,8 +494,10 @@ void foo(void);
|
||||
\end{lstlisting}
|
||||
|
||||
|
||||
The other use of \keyword{void} is when you are defining an \keyword{lvalue}.
|
||||
A \keyword{void *} pointer is just a memory address. It is specified as an incomplete type meaning that you cannot dereference it but it can be promoted to any time to any other type. Pointer arithmetic with this pointer is undefined behavior.
|
||||
The other use of \keyword{void} is the generic pointer type \keyword{void *}.
|
||||
A \keyword{void *} pointer is just a memory address.
|
||||
\keyword{void} is an incomplete type, meaning that you cannot dereference a \keyword{void *}, but it can be converted to any other object pointer type without a cast.
|
||||
Standard C does not allow pointer arithmetic on a \keyword{void *}, although gcc and clang permit it as an extension (see the Pointers section).
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
int *array = void_ptr; // No cast needed
|
||||
|
||||
+1
-1
@@ -91,7 +91,7 @@ If this sounds familiar, it is what C++ originally intended to do before the sta
|
||||
|
||||
\subsection{Pointer Arithmetic}
|
||||
|
||||
In addition to adding to an integer, pointers can be added to.
|
||||
Just as you can add an integer to an integer, you can add an integer to a pointer.
|
||||
However, the pointer type is used to determine how much to increment the pointer.
|
||||
A pointer is moved over by the value added times the size of the underlying type.
|
||||
For char pointers, this is trivial because characters are always one byte.
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
\epigraph{To thy happy children of the future, those of the past send greetings.}{Alma Mater}
|
||||
|
||||
At the University of Illinois at Urbana-Champaign, we fundamentally believe that we have a right to make the university better for all future students.
|
||||
It is a message etched into our Alma Mater and makes up the DNA of our course staff.
|
||||
That belief is a message etched into our Alma Mater and makes up the DNA of our course staff.
|
||||
As such, we created the coursebook.
|
||||
The coursebook is a free and open systems programming textbook that anyone can read, contribute to, and modify for now and forever.
|
||||
We don't think information should be behind a walled garden, and we truly believe that complex concepts can be explained simply and fully, for anyone to understand.
|
||||
@@ -11,10 +11,10 @@ The goal of this book is to teach you the basics and give you some intuition int
|
||||
|
||||
Like any good book, it isn't complete.
|
||||
We still have plenty of examples, ideas, typos, and chapters to work on.
|
||||
If you find any issues, please file an \href{https://github.com/illinois-cs241/coursebook/issues}{issue} or email a list of typos to \href{http://cs341.cs.illinois.edu/staff}{CS 341 course staff}, and we'll be happy to work on it.
|
||||
If you find any issues, please file an \href{https://github.com/illinois-cs241/coursebook/issues}{issue} or email a list of typos to \href{https://cs341.cs.illinois.edu/staff}{CS 341 course staff}, and we'll be happy to work on it.
|
||||
We are constantly trying to make the book better for students a year and ten years from now.
|
||||
|
||||
This work is based on the original coursebook located \href{https://github.com/angrave/SystemProgramming/wiki}{at this url}.
|
||||
This work is based on the original \href{https://github.com/angrave/SystemProgramming/wiki}{SystemProgramming wiki coursebook}.
|
||||
All these people's hard work is included in the section below.
|
||||
|
||||
Oh and the duck? Keep reading until synchronization :).
|
||||
@@ -25,5 +25,8 @@ Thanks again and happy reading!
|
||||
|
||||
\section{Authors}
|
||||
|
||||
\lstinputlisting{AUTHORS.md}
|
||||
% AUTHORS.md stays the single source of truth. \lstinputlisting (rather than a
|
||||
% TeX read loop) keeps it in the EPUB and wiki too, since pandoc reads it; the
|
||||
% options only restyle the PDF so the list typesets as plain text, not code.
|
||||
\lstinputlisting[basicstyle=\small\rmfamily,frame=none,backgroundcolor=\color{white},columns=fullflexible,keepspaces=true,breaklines=true,numbers=none,aboveskip=0pt]{AUTHORS.md}
|
||||
|
||||
|
||||
+5
-5
@@ -219,13 +219,13 @@ We will assume that this is for a single-level page table.
|
||||
\item If the translation fails, report an invalid address
|
||||
\item Otherwise,
|
||||
\begin{enumerate}
|
||||
\item If the TLB contains the physical memory, get the physical frame from the TLB and perform the read and write.
|
||||
\item If the TLB contains the physical memory, get the physical frame from the TLB and perform the read or write.
|
||||
\item If the page exists in memory, check if the process has permissions
|
||||
to perform the operation on the page meaning the process has access
|
||||
to the page, and it is reading from the page/writing to a page
|
||||
that it has permission to do so.
|
||||
\begin{enumerate}
|
||||
\item If so then do the dereference provide the address, cache the results in the TLB
|
||||
\item If so, translate the address to the physical frame, perform the read or write, and cache the translation in the TLB.
|
||||
\item Otherwise, trigger a hardware interrupt. The kernel
|
||||
will most likely send a SIGSEGV or a Segmentation Violation.
|
||||
\end{enumerate}
|
||||
@@ -250,7 +250,7 @@ Read-only frames can then be safely shared between multiple processes.
|
||||
For example, the C-library instruction code can be shared between all processes that dynamically load the code into the process memory.
|
||||
Each process can only read that memory.
|
||||
Meaning that if a program tries to write to a read-only page in memory, it will \keyword{SEGFAULT}.
|
||||
That is why sometimes memory accesses SEGFAULT and sometimes they don't, it all depends on if your hardware says that a program can access.
|
||||
That is why sometimes memory accesses SEGFAULT and sometimes they don't, it all depends on whether the hardware says that the program can access that page in that way.
|
||||
|
||||
Also, processes can share a page with a child process using the \keyword{mmap} system call.
|
||||
\keyword{mmap} is an interesting call because instead of tying each virtual address to a physical frame, it ties it to something else. It is an important distinction that we are talking about mmap and not memory-mapped IO in general. The \keyword{mmap} system call can't reliably be used to do other memory-mapped operations like communicate with GPUs and write pixels to the screen -- this is mainly hardware dependent.
|
||||
@@ -648,7 +648,7 @@ What happens when a process tries to write when there are no readers left?
|
||||
\end{verbatim}
|
||||
|
||||
Tip: Notice only the writer (not a reader) can use this signal.
|
||||
To inform the reader that a writer is closing their end of the pipe, a program could write your special byte (e.g.~0xff) or a message (\keyword{"Bye!"}).
|
||||
To inform the reader that a writer is closing their end of the pipe, a program could write a special byte (e.g.~0xff) or a message (\keyword{"Bye!"}).
|
||||
|
||||
Here's an example of catching this signal that fails! Can you see why?
|
||||
|
||||
@@ -931,7 +931,7 @@ On Linux, there are two abstractions with files. The first is the Linux \keyword
|
||||
\item \keyword{open} takes a path to a file and creates a file descriptor entry in the process table. If the file is inaccessible, it errors out.
|
||||
\item \keyword{read} takes a certain number of bytes that the kernel has received and reads them into a user-space buffer. If the file is not open in read mode, this will break.
|
||||
\item \keyword{write} outputs a certain number of bytes to a file descriptor. If the file is not open in write mode, this will break. This may be buffered internally.
|
||||
\item \keyword{close} removes a file descriptor from a process' file descriptors. This always succeeds for a valid file descriptor.
|
||||
\item \keyword{close} removes a file descriptor from a process's file descriptors. This always succeeds for a valid file descriptor.
|
||||
\item \keyword{lseek} takes a file descriptor and moves it to a certain position. It can fail if the seek is out of bounds.
|
||||
\item \keyword{fcntl} is the catch-all function for file descriptors. Set file locks, read, write, edit permissions, etc.
|
||||
\end{itemize}
|
||||
|
||||
+3
-3
@@ -6,7 +6,7 @@
|
||||
|
||||
Memory allocation is important!
|
||||
Allocating and deallocating heap memory is one of the most common operations in any application.
|
||||
The heap at the system level is a contiguous series of addresses that the program can expand or contract and use as its accord \cite{mallocinternals}.
|
||||
The heap at the system level is a contiguous series of addresses that the program can expand or contract and use as it sees fit \cite{mallocinternals}.
|
||||
In POSIX, this is called the system break.
|
||||
We use \keyword{sbrk} to move the system break.
|
||||
Most programs don't interact directly with this call, they use a memory allocation system around it to handle chunking up and keeping track of which memory is allocated and which is freed.
|
||||
@@ -210,7 +210,7 @@ Or it could split one of the other two free holes.
|
||||
These choices represent different placement strategies.
|
||||
Whichever hole is chosen, the allocator will need to split the hole into two.
|
||||
The first is the newly allocated space, which will be returned to the program, and the second is a smaller hole if there is spare space left over.
|
||||
A perfect-fit strategy finds the smallest hole that is of sufficient size (at least 2KiB):
|
||||
A best fit strategy finds the smallest hole that is of sufficient size (at least 2KiB):
|
||||
|
||||
\begin{figure}[H]
|
||||
\centering
|
||||
@@ -495,7 +495,7 @@ With the above description, it's possible to build a memory allocator.
|
||||
Its main advantage is simplicity - at least simple compared to other allocators!
|
||||
Allocating memory is a worst-case linear time operation -- search linked lists for a sufficiently large free block.
|
||||
De-allocation is constant time.
|
||||
No more than 3 blocks will need to coalesce into a single block, and using a most recently used block scheme only one linked list entry.
|
||||
No more than 3 blocks will need to coalesce into a single block, and using a most recently used block scheme, only one linked list entry needs to be updated.
|
||||
|
||||
Using this allocator it is possible to experiment with different placement strategies.
|
||||
For example, the allocator could start searching from the last deallocated block.
|
||||
|
||||
+17
-13
@@ -67,13 +67,15 @@ IPv4 was designed at a time when the idea of 4 billion devices connected to the
|
||||
IPv4 addresses are written typically in a sequence of four octets delimited by periods "255.255.255.0" for example.
|
||||
|
||||
Each IPv4 datagram includes a small header - typically 20 octets, that includes a source and destination address.
|
||||
Conceptually the source and destination addresses can be split into two: a network number the upper bits and lower bits represent a particular host number on that network.
|
||||
Conceptually, each source and destination address can be split into two parts: the upper bits are a network number, and the lower bits represent a particular host number on that network.
|
||||
|
||||
A newer packet protocol IPv6 solves many of the limitations of IPv4 like making routing tables simpler and 128-bit addresses.
|
||||
Adoption was slow at first: in 2018, only a small fraction of web traffic used IPv6 \cite{internet_society_2018}.
|
||||
By the mid-2020s, Google measured roughly half of its user traffic arriving over IPv6 \cite{google_ipv6_stats}.
|
||||
We write IPv6 addresses in a sequence of eight, four hexadecimal delimiters like "1F45:0000:0000:0000:0000:0000:0000:0000".
|
||||
Since that can get unruly, we can omit the zeros "1F45::". A machine can have an IPv6 address and an IPv4 address.
|
||||
We write IPv6 addresses as eight groups of four hexadecimal digits separated by colons, like "1F45:0000:0000:0000:0000:0000:0000:0000".
|
||||
Since that can get unruly, we can replace one run of consecutive all-zero groups with "::", so the address above becomes "1F45::".
|
||||
The "::" may appear only once in an address; otherwise, you couldn't tell how many zero groups each one stands for.
|
||||
A machine can have an IPv6 address and an IPv4 address.
|
||||
|
||||
There are special IP Addresses.
|
||||
One such in IPv4 is \keyword{127.0.0.1}, IPv6 as \keyword{0:0:0:0:0:0:0:1} or \keyword{::1} also known as localhost.
|
||||
@@ -247,7 +249,7 @@ The 16-bit source and destination ports come first, then the 32-bit sequence and
|
||||
Most services on the Internet today use TCP because it efficiently hides the complexity of the lower, packet-level nature of the Internet.
|
||||
TCP or Transmission Control Protocol is a connection-based protocol that is built on top of IPv4 and IPv6 and therefore can be described as ``TCP/IP'' or ``TCP over IP''.
|
||||
TCP creates a \emph{pipe} between two machines and abstracts away the low-level packet-nature of the Internet. Thus, under most conditions, bytes sent over a TCP connection are delivered and uncorrupted.
|
||||
High performance and error-prone code won't even assume that!
|
||||
High-performance, error-tolerant code won't even assume that bytes are delivered!
|
||||
|
||||
TCP has many features that set it apart from the other transport protocol UDP.
|
||||
|
||||
@@ -255,8 +257,8 @@ TCP has many features that set it apart from the other transport protocol UDP.
|
||||
\item Ports
|
||||
With IP, you are only allowed to send packets to a machine.
|
||||
If you want one machine to handle multiple flows of data, you have to do it manually with IP.
|
||||
TCP gives the programmer a set of virtual sockets.
|
||||
Clients specify the socket that you want the packet sent to and the TCP protocol makes sure that applications that are waiting for packets on that port receive that.
|
||||
TCP gives the programmer a set of virtual ports.
|
||||
Clients specify the port that they want the packet sent to and the TCP protocol makes sure that applications that are waiting for packets on that port receive that.
|
||||
A process can listen for incoming packets on a particular port.
|
||||
However, only processes with super-user (root) access can listen on ports less than 1024.
|
||||
Any process can listen on ports 1024 or higher.
|
||||
@@ -502,7 +504,7 @@ freeaddrinfo(result);
|
||||
The next piece of code sends the request. Here is what each header means.
|
||||
|
||||
\begin{enumerate}
|
||||
\item "GET \%s HTTP/1.0" This is the request verb interpolated with the path. This means to perform the GET verb on the path using the HTTP/1.0 method.
|
||||
\item "GET \%s HTTP/1.0" This is the request verb interpolated with the path. This means to perform the GET verb on the path using the HTTP/1.0 protocol version.
|
||||
\item "Connection: close" Means that as soon as the request is over, please close the connection. This line won't be used for any other connections.
|
||||
This is a little redundant given that HTTP 1.0 doesn't allow you to send multiple requests, but it is better to be explicit given there are non-conformant technologies.
|
||||
\item "Accept: */*" This means that the client is willing to accept anything.
|
||||
@@ -561,7 +563,8 @@ In general, there are six parts:
|
||||
\item The protocol ``HTTP/1.0''
|
||||
\item A new line (\keyword{\textbackslash r\textbackslash n}). Requests always have a carriage return.
|
||||
\item Any other knobs or switch parameters
|
||||
\item The actual body of the request delimited by two new lines. The body of the request is either if the size is specified or until the receiver closes their connection.
|
||||
\item The actual body of the request, which follows a blank line (two new lines in a row).
|
||||
The receiver reads the body either for the number of bytes given in the \keyword{Content-Length} header or, if no size is specified, until the sender closes the connection.
|
||||
\end{enumerate}
|
||||
|
||||
The server's first response line describes the HTTP version used and whether the request is successful using a 3 digit response code.
|
||||
@@ -829,7 +832,7 @@ A streaming video signal may send picture updates using UDP.
|
||||
\item Manual Flow/Congestion Control
|
||||
You have to manually manage the flow and congestion control which is a double-edged sword.
|
||||
On one hand, you have full control over everything.
|
||||
On the other hand, TCP has \textit{decades} of optimization, meaning your protocol for its use cases needs to be more efficient that to be more beneficial to use it.
|
||||
On the other hand, TCP has \textit{decades} of optimization, so a custom protocol over UDP is only worth it if it beats TCP for your particular use case.
|
||||
\item Multicast
|
||||
This is one thing that you can only do with UDP.
|
||||
This means that you can send a message to every peer connected to a particular router that is part of a particular group.
|
||||
@@ -905,7 +908,7 @@ There is no idea of if the packet arrives, is processed, etc.
|
||||
|
||||
\subsection{UDP Server}
|
||||
|
||||
There are a variety of function calls available to send UDP sockets.
|
||||
There are a variety of function calls available to send data over UDP sockets.
|
||||
We will use the newer \keyword{getaddrinfo} to help set up a socket structure.
|
||||
Remember that UDP is a simple packet-based (`datagram') protocol.
|
||||
There is no connection to set up between the two hosts.
|
||||
@@ -1358,7 +1361,7 @@ int getHighScore(char* game) {
|
||||
Using a string format may be a little inefficient.
|
||||
A good example of this marshaling is Golang's gRPC or Google RPC. There is a version in C as well if you want to check that out.
|
||||
|
||||
The server stub code will receive the request, unmarshal the request into a valid in-memory data call the underlying implementation and send the result back to the caller.
|
||||
The server stub code will receive the request, unmarshal the request into a valid in-memory representation, call the underlying implementation, and send the result back to the caller.
|
||||
Often the underlying library will do this for you.
|
||||
|
||||
To implement RPC you need to decide and document which conventions you will use to serialize the data into a byte sequence.
|
||||
@@ -1375,7 +1378,7 @@ To marshal a struct, decide which fields need to be serialized.
|
||||
It may be unnecessary to send all data items.
|
||||
For example, some items may be irrelevant to the specific RPC or can be re-computed by the server from the other data items present.
|
||||
|
||||
To marshal a linked list, it is unnecessary to send the link pointers, stream the values.
|
||||
To marshal a linked list, it is unnecessary to send the link pointers; instead, stream the values.
|
||||
As part of unmarshaling, the server can recreate a linked list structure from the byte sequence.
|
||||
|
||||
By starting at the head node/vertex, a simple tree can be recursively visited to create a serialized version of the data.
|
||||
@@ -1383,7 +1386,8 @@ A cyclic graph will usually require additional memory to ensure that each edge a
|
||||
|
||||
\subsection{Interface Description Language}
|
||||
|
||||
Writing stub code by hand is painful, tedious, error-prone, difficult to maintain and difficult to reverse engineer the wire protocol from the implemented code.
|
||||
Writing stub code by hand is painful, tedious, error-prone, and difficult to maintain.
|
||||
It is also difficult to reverse engineer the wire protocol from the implemented code.
|
||||
A better approach is to specify the data objects, messages, and services to automatically generate the client and server code.
|
||||
A modern example of an Interface Description Language is Google's Protocol Buffer .proto files.
|
||||
|
||||
|
||||
@@ -11,6 +11,7 @@ Sit back and scroll through this chapter as we tell you about the problems of pa
|
||||
Even if you are dealing with something much higher level like web-development, everything relates back to the system.
|
||||
|
||||
\section{Shell Shock}
|
||||
\label{pm:shellshock}
|
||||
|
||||
Required: Appendix/Shell
|
||||
|
||||
@@ -32,6 +33,7 @@ Also, you can harden your system to never perform exec calls to perform tasks (i
|
||||
Although you don't have flexibility, you have peace of mind about what you allow users to do.
|
||||
|
||||
\section{Heartbleed}
|
||||
\label{pm:heartbleed}
|
||||
|
||||
Required: Intro to C
|
||||
|
||||
@@ -45,6 +47,7 @@ Lessons Learned: Check your buffers!
|
||||
Know the difference between a buffer and a string.
|
||||
|
||||
\section{Dirty Cow}
|
||||
\label{pm:dirtycow}
|
||||
|
||||
Required: Processes/Virtual Memory
|
||||
|
||||
@@ -58,28 +61,42 @@ This can be done to the effective user id bit and the process can pretend it was
|
||||
Lessons Learned: Spinlocks in the kernel are hard.
|
||||
|
||||
\section{Meltdown}
|
||||
\label{pm:meltdown}
|
||||
|
||||
There is an example of this in the background section.
|
||||
Meltdown is a close cousin of Spectre: both abuse out-of-order execution to leak data the program shouldn't be able to read.
|
||||
See Section~\ref{sec:spectre} in the security chapter.
|
||||
|
||||
\section{Spectre}
|
||||
\label{pm:spectre}
|
||||
|
||||
Check in the security section.
|
||||
See Section~\ref{sec:spectre} in the security chapter.
|
||||
|
||||
\section{Mars Pathfinder}
|
||||
\label{pm:pathfinder}
|
||||
|
||||
Required sections: Synchronization and a bit of Scheduling
|
||||
|
||||
\href{https://www.microsoft.com/en-us/research/people/mbj/#!just-for-fun}{Pathfinder Link}
|
||||
|
||||
The Mars Pathfinder was a mission that tried to collect climate data on Mars. The finder uses a single bus to communicate with different parts. Since this was 1997, the hardware itself didn't have advanced features like efficient locking so it was up to the operating system developers to regulate that with mutexes. The architecture was pretty simple. There was a thread that controlled data along the information bus, communications thread, and data collection thread with high, regular, and low priorities with respect to scheduling. The other caveat is that if an interrupt happened at some interval and a task is running and a task is to be scheduled, the task that has the higher priority wins.
|
||||
The Mars Pathfinder was a 1997 mission that, among other things, collected weather data on Mars.
|
||||
The lander used a single ``information bus'' to pass data between its different parts, and access to the bus was protected by a mutex.
|
||||
The architecture was pretty simple.
|
||||
There was a bus management thread with high priority, a communications thread with medium priority, and a meteorological data collection thread with low priority.
|
||||
The scheduler was preemptive: whenever a higher-priority task became ready to run, it took the CPU from any lower-priority task.
|
||||
|
||||
The pattern that caused everything to start failing was the data collection thread starts writing to the bus, the information bus thread is waiting on the data.
|
||||
Then the communication thread comes in to preempt the other lower priority thread \textbf{while the lower priority thread still held the mutex}. This means when the regular priority thread tried to lock the bus, the rover would deadlock.
|
||||
After some time the system would reset, but that isn't good to leave to chance.
|
||||
The pattern that caused everything to start failing went like this.
|
||||
The low-priority data collection thread locked the mutex to publish its data on the bus.
|
||||
The high-priority bus thread then tried to lock the same mutex and had to wait.
|
||||
Then the medium-priority communications thread became ready and preempted the low-priority thread \textbf{while the low-priority thread still held the mutex}.
|
||||
The communications thread ran for a long time, so the low-priority thread couldn't finish and release the mutex, and the high-priority bus thread stayed blocked behind both of them.
|
||||
This is called \keyword{priority inversion}: a medium-priority task effectively kept a high-priority task from running.
|
||||
After a while, a watchdog timer noticed that the bus thread hadn't run and reset the whole system, losing data each time.
|
||||
The fix, uploaded to the spacecraft, was to turn on \keyword{priority inheritance} for that mutex, so a low-priority thread holding the mutex temporarily runs at the priority of the highest-priority thread waiting for it.
|
||||
|
||||
Moral of the lesson? Don't have the applications themselves deal with the synchronization. Define a module that handles mutex locking and have the module communicate through files, IPC, etc.
|
||||
|
||||
\section{Mars Again}
|
||||
\label{pm:mars_memory}
|
||||
|
||||
Required Sections: Malloc
|
||||
|
||||
@@ -88,6 +105,7 @@ Required Sections: Malloc
|
||||
The short of it is that they ran out of memory. The long of it is that they ran out of memory, disk space, and swap space. The moral of the story? Make sure to write code that can handle file failures and can handle files when they close and go out of memory, so the operating system can hot swap files to free up memory. Also clean up files, assume that your temp directory is roughly a hundredth or a thousandth of the total size and use that.
|
||||
|
||||
\section{Year 2038}
|
||||
\label{pm:y2038}
|
||||
|
||||
Required sections: Intro to C
|
||||
|
||||
@@ -103,6 +121,7 @@ Stay tuned to see what happens.
|
||||
Lessons learned: Plan like your application will be huge one day.
|
||||
|
||||
\section{Northeast Blackout of 2003}
|
||||
\label{pm:blackout2003}
|
||||
|
||||
Required Sections: Synchronization
|
||||
|
||||
@@ -115,6 +134,7 @@ Lessons Learned: Modularize your code to localize failures (i.e. keeping race co
|
||||
|
||||
|
||||
\section{Apple IOS Unicode Handling}
|
||||
\label{pm:ios_unicode}
|
||||
|
||||
Required Sections: Intro to C
|
||||
|
||||
@@ -128,6 +148,7 @@ Undefined behavior means anything though, and a lot of varied things did happen
|
||||
Lessons Learned: Fuzz your kernel.
|
||||
|
||||
\section{Apple SSL Verification}
|
||||
\label{pm:apple_ssl}
|
||||
|
||||
Required Sections: Intro to C
|
||||
|
||||
@@ -140,6 +161,7 @@ Naturally, hackers were able to get away with some pretty crazy site names.
|
||||
Lessons Learned: Always bracket if statements, use gotos sparingly. Chances are if you need to use a goto, write another function or a switch statement with fall throughs (still bad).
|
||||
|
||||
\section{Sony Rootkit Installation}
|
||||
\label{pm:sony_rootkit}
|
||||
|
||||
Required Sections: Intro to C/Processes
|
||||
|
||||
@@ -153,13 +175,14 @@ With 22 million Music CDs, they required users to install a rootkit on their ope
|
||||
|
||||
Privacy concerns aside, and believe me there are a lot of them, the big problem was that this rootkit was a backdoor for everyone's systems if programmed incorrectly.
|
||||
A rootkit is a piece of code usually installed kernel-side that keeps track of almost anything that a user does.
|
||||
What websites visited, what clicks or keys typed etc.
|
||||
It can see which websites are visited, which buttons are clicked, which keys are typed, etc.
|
||||
If a hacker finds out about this and there is a way to access that API from the user space level, that means any program can find out important information about your device.
|
||||
Needless to say, people were angry.
|
||||
|
||||
Lessons Learned: Get an antivirus and/or apparmor and make sure that an application is only requesting permissions that make sense. If you are torn, try something like Windows sandbox or keep a Sacrificial VM around to see if installing it makes your computer horrible. Don't trust certificates, trust code.
|
||||
|
||||
\section{The Woes of Shell Scripting}
|
||||
\label{pm:steam_rm}
|
||||
|
||||
Required Sections: Appendix/Shell
|
||||
|
||||
@@ -179,19 +202,22 @@ What happens if that directory doesn't exist, for example because the user moved
|
||||
Lessons Learned: Do parameter checks, always always always \keyword{set -e} on a script and if you expect a command to fail, explicitly list it. You can also alias rm to mv and then delete the trash later.
|
||||
|
||||
\section{Appnexus Double Free}
|
||||
\label{pm:appnexus}
|
||||
|
||||
Required Sections: Intro to C/Malloc
|
||||
|
||||
\href{https://techblog.appnexus.com/2013-09-17-outage-postmortem-586b19ae4307}{Double Free}
|
||||
|
||||
Appnexus uses an asynchronous garbage collector that reclaims different parts of the heap when it believes that objects are unused.
|
||||
The architecture is that an element is in the unavailable list and then it is taken out to a to-be-freed list.
|
||||
After a certain time if that element was unused, it is freed and added to the free list.
|
||||
This is fine until two threads try to delete the same object at once, adding to the list twice. After less time, one of the objects was deleted, the delete was announced to other computers.
|
||||
To keep latency low, when an in-memory object is deleted, it is unlinked from other objects and its memory is scheduled to be freed at a safe time in the future, when no thread could still be using it.
|
||||
In 2013, a rarely-changed object was deleted by a data update, and a bug in the code deleted that object twice.
|
||||
The update passed validation because the actual free hadn't happened yet, so it was distributed to all of the roughly 900 ad-serving servers.
|
||||
When the scheduled frees finally ran, the double free crashed all of them at nearly the same time.
|
||||
|
||||
Lessons Learned: Avoid making hacky software if you need to. Modularize, and set memory limits, and monitor different parts of your code and optimize by hand. There is no general catch-all garbage collector that fits everyone. Even highly tested ones like the JVM need some nudges if you want to get performance out of them.
|
||||
|
||||
\section{ATT Cascading Failures - 1990}
|
||||
\label{pm:att1990}
|
||||
|
||||
Required Sections: Intro to C
|
||||
|
||||
@@ -199,9 +225,10 @@ Required Sections: Intro to C
|
||||
|
||||
The bug is explained well at the link above.
|
||||
We recommend reading to learn more.
|
||||
A series of network delays that caused some telephone switches across the country to think that other switches were operable when they weren't.
|
||||
When the switches came back online, they realized they had a huge backlog of calls to route and began doing so.
|
||||
Other routing failures and restarts only compounded the problem.
|
||||
A switch in New York reset itself after a fault and told its neighbors that it was out of service.
|
||||
When it came back online, it began routing the backlog of calls that had built up and sent a burst of closely timed messages to other switches.
|
||||
Because of a misplaced \keyword{break} statement in the C code, a switch that received a second message while still processing the first overwrote its own data and reset itself.
|
||||
When each of those switches came back online, it sent its own burst of messages, so the resets cascaded across all 114 switches in the network for about nine hours.
|
||||
|
||||
Lessons Learned: Not using C would've actually helped here because of more rigorous fuzzing (though C++ in this day and age would be worse with its language constructs).
|
||||
The real moral of the story is networks are random and expect any jump at any point in your code.
|
||||
|
||||
+12
-12
@@ -7,7 +7,7 @@ An operating system is a program that provides an interface between hardware and
|
||||
The operating system manages hardware and gives user programs a uniform way of interacting with hardware as long as the operating system can be installed on that hardware.
|
||||
Although this idea sounds like it is the end-all, we know that there are many different operating systems with their own quirks and standards.
|
||||
As a solution to that, there is another layer of abstraction: POSIX or portable operating systems interface.
|
||||
This is a standard (or many standards now) that an operating system must implement to be POSIX compatible -- most systems that we'll be studying are almost POSIX compatible due more to political reasons.
|
||||
This is a standard (or many standards now) that an operating system must implement to be POSIX compatible -- most systems that we'll be studying are ``mostly POSIX-compliant'' rather than officially certified, because formal certification is costly and many vendors don't bother.
|
||||
|
||||
Before we talk about POSIX systems, we should understand what the idea of a kernel is generally.
|
||||
In an operating system (OS), there are two spaces: kernel space and user space.
|
||||
@@ -98,7 +98,7 @@ printf("%d\n", secrets);
|
||||
\end{lstlisting}
|
||||
|
||||
On two different terminals, they would both print out 1 not 2.
|
||||
Even if we changed the code to attempt to affect other process instances, there would be no way to change another process' state unintentionally.
|
||||
Even if we changed the code to attempt to affect other process instances, there would be no way to change another process's state unintentionally.
|
||||
However, there are other intentional ways to change the program states of other processes.
|
||||
|
||||
\section{Process Contents}
|
||||
@@ -139,7 +139,7 @@ Furthermore, the initialized data segment is divided into a readable and writabl
|
||||
\item \textbf{Initialized Data Segment}
|
||||
This contains all of a program's globals and any other static variables.
|
||||
|
||||
This section starts at the end of the text segment and starts at a constant size because the number of globals is known at compile time. The end of the data segment is called the \keyword{program break} and can be extended via the use of brk / sbrk.
|
||||
This section starts at the end of the text segment and has a constant size because the number of globals is known at compile time. The end of the data segment is called the \keyword{program break} and can be extended via the use of brk / sbrk.
|
||||
|
||||
This section is writable \cite[P. 124]{van1994expert}.
|
||||
Most notably, this section contains variables that were initialized with a static initializer, as follows:
|
||||
@@ -455,7 +455,7 @@ Here is a summary of what is relevant:
|
||||
\item Since we have copy on write (COW), read-only memory addresses are shared between processes.
|
||||
\item If a program sets up certain regions of memory, they can be shared between processes.
|
||||
\item Signal handlers are inherited but can be changed.
|
||||
\item The process' current working directory (often abbreviated to CWD) is inherited but can be changed.
|
||||
\item The process's current working directory (often abbreviated to CWD) is inherited but can be changed.
|
||||
\item Environment variables are inherited but can be changed.
|
||||
\end{enumerate}
|
||||
|
||||
@@ -655,7 +655,7 @@ A process can only have 256 return values, the rest of the bits are informationa
|
||||
However, the kernel has an internal way of keeping track of signaled, exited, or stopped processes.
|
||||
This API is abstracted so that the kernel developers are free to change it at will.
|
||||
Remember: these macros only make sense if the precondition is met.
|
||||
For example, a process' exit status (\keyword{WEXITSTATUS}) is only defined if the process exited normally (\keyword{WIFEXITED}), and the signal that killed it (\keyword{WTERMSIG}) is only defined if the process was signaled (\keyword{WIFSIGNALED}).
|
||||
For example, a process's exit status (\keyword{WEXITSTATUS}) is only defined if the process exited normally (\keyword{WIFEXITED}), and the signal that killed it (\keyword{WTERMSIG}) is only defined if the process was signaled (\keyword{WIFSIGNALED}).
|
||||
The macros will not do the checking for the program, so it's up to the programmer to make sure the logic is correct.
|
||||
As an example above, the program should use the \keyword{WIFSTOPPED} to check if a process was stopped and then the \keyword{WSTOPSIG} to find the signal that stopped it.
|
||||
As such, there is no need to memorize the following. This is a high-level overview of how information is stored inside the status variables. From the \keyword{sys/wait.h} of an old Berkeley Standard Distribution (BSD) kernel \cite{sys/wait.h}:
|
||||
@@ -681,15 +681,15 @@ Usually, UNIX programs are not designed to follow this policy, for the sake of s
|
||||
|
||||
\subsection{Zombies and Orphans}
|
||||
|
||||
It is good practice to wait on your process' children.
|
||||
If a parent doesn't wait on your children they become what are called zombies.
|
||||
It is good practice for a process to wait on its children.
|
||||
If a parent doesn't wait on its children, they become what are called zombies.
|
||||
Zombies are created when a child terminates and then takes up a spot in the kernel process table for your process.
|
||||
The process table keeps track of the following information about a process: PID, status, and how it was killed.
|
||||
The only way to get rid of a zombie is to wait on your children.
|
||||
If a long-running parent never waits for your children, it may lose the ability to fork.
|
||||
The only way to get rid of a zombie is for its parent to wait on it.
|
||||
If a long-running parent never waits for its children, it may lose the ability to fork.
|
||||
|
||||
Having said that, a program doesn't always need to wait for your children!
|
||||
Your parent process can continue to execute code without having to wait for the child process.
|
||||
Having said that, a program doesn't always need to wait for its children!
|
||||
A parent process can continue to execute code without having to wait for the child process.
|
||||
If a parent dies without waiting on its children, a process can orphan its children.
|
||||
Once a parent process completes, any of its children will be assigned to \keyword{init} - the first process, whose PID is 1.
|
||||
Therefore, these children would see \keyword{getppid()} return a value of 1.
|
||||
@@ -1008,7 +1008,7 @@ Note that we aren't expecting you to memorize the man page.
|
||||
\item
|
||||
My terminal is anchored to PID = 1337 and has become unresponsive. Write me the terminal command and the C code to send \keyword{SIGQUIT} to it.
|
||||
\item
|
||||
Can one process alter another process' memory through normal means? Why?
|
||||
Can one process alter another process's memory through normal means? Why?
|
||||
\item
|
||||
Where is the heap, stack, data, and text segment? Which segments can a program write to? What are invalid memory addresses?
|
||||
\item
|
||||
|
||||
+5
-6
@@ -198,13 +198,12 @@ char* next_ticket() {
|
||||
\item What is a free list?
|
||||
\item What are some different ways of inserting into a free list?
|
||||
\item What are the benefits and drawbacks to first fit, worst fit, best fit?
|
||||
\item When would a trivial malloc implementation
|
||||
\item When would the following trivial malloc implementation be acceptable?
|
||||
\begin{lstlisting}[language=C]
|
||||
void *malloc(int size) {
|
||||
return (void *)sbrk(size);
|
||||
}
|
||||
\end{lstlisting}
|
||||
be acceptable?
|
||||
\end{enumerate}
|
||||
|
||||
\section{Threading and Synchronization}
|
||||
@@ -514,13 +513,13 @@ void xout(char* filename) {
|
||||
\item What is the sticky bit?
|
||||
\item What is a virtual file system?
|
||||
\item What is RAID?
|
||||
\item In an \keyword{ext2} filesystem how many inodes are read from disk to access the first byte of the file \keyword{/dir1/subdirA/notes.txt} ? Assume the directory names and inode numbers in the root directory (but not the inodes themselves) are already in memory.
|
||||
\item In an \keyword{ext2} filesystem how many inodes are read from disk to access the first byte of the file \keyword{/dir1/subdirA/notes.txt}? Assume the directory names and inode numbers in the root directory (but not the inodes themselves) are already in memory.
|
||||
|
||||
\item In an \keyword{ext2} filesystem what is the minimum number of disk blocks that must be read from disk to access the first byte of the file \keyword{/dir1/subdirA/notes.txt} ? Assume the directory names and inode numbers in the root directory and all inodes are already in memory.
|
||||
\item In an \keyword{ext2} filesystem what is the minimum number of disk blocks that must be read from disk to access the first byte of the file \keyword{/dir1/subdirA/notes.txt}? Assume the directory names and inode numbers in the root directory and all inodes are already in memory.
|
||||
|
||||
\item In an \keyword{ext2} filesystem with 32 bit addresses and 4KiB disk blocks, an inode can store 10 direct disk block numbers. What is the minimum file size required to require a single indirection table? ii) a double indirection table?
|
||||
|
||||
\item Fix the shell command \keyword{chmod} below to set the permission of a file \keyword{secret.txt} so that the owner can read, write, and execute permissions the group can read and everyone else has no access.
|
||||
\item Fix the shell command \keyword{chmod} below to set the permission of a file \keyword{secret.txt} so that the owner can read, write, and execute it, the group can only read it, and everyone else has no access.
|
||||
|
||||
\begin{lstlisting}[language=bash]
|
||||
$ chmod 000 secret.txt
|
||||
@@ -586,7 +585,7 @@ $ chmod 000 secret.txt
|
||||
|
||||
\item When would you call bind on a TCP client?
|
||||
|
||||
\item What is the purpose of socket bind listen accept ?
|
||||
\item What is the purpose of \keyword{socket}, \keyword{bind}, \keyword{listen}, and \keyword{accept}?
|
||||
|
||||
\item Which of the above calls can block, waiting for a new client to connect?
|
||||
|
||||
|
||||
@@ -54,7 +54,7 @@ The throughput might be measured by a system value, for example, the I/O through
|
||||
The latency might be measured by the response time -- elapsed time before a process can start to send a response -- or wait time or turnaround time -- the elapsed time to complete a task.
|
||||
Different schedulers offer different optimization trade-offs that may be appropriate for desired use.
|
||||
There is no optimal scheduler for all possible environments and goals.
|
||||
For example, Shortest Job First will minimize total wait time across all jobs but in interactive (UI) environments it would be preferable to minimize response time at the expense of some throughput, while FCFS seems intuitively fair and easy to implement but suffers from the Convoy Effect.
|
||||
For example, Shortest Job First will minimize total wait time across all jobs but in interactive (UI) environments it would be preferable to minimize response time at the expense of some throughput, while FCFS seems intuitively fair and easy to implement but suffers from the convoy effect.
|
||||
Arrival time is the time at which a process first arrives at the ready queue, and is ready to start executing.
|
||||
If a CPU is idle, the arrival time would also be the starting time of execution.
|
||||
|
||||
@@ -63,7 +63,7 @@ If a CPU is idle, the arrival time would also be the starting time of execution.
|
||||
Without preemption, processes will run until they are unable to utilize the CPU any further.
|
||||
For example the following conditions would remove a process from the CPU and the CPU would be available to be scheduled for other processes.
|
||||
The process terminates due to a signal, is blocked waiting for a concurrency primitive, or exits normally.
|
||||
Thus once a process is scheduled it will continue even if another process with a high priority appears on the ready queue.
|
||||
Thus once a process is scheduled it will continue even if another process with a higher priority appears on the ready queue.
|
||||
|
||||
With preemption, the existing processes may be removed immediately if a more preferred process is added to the ready queue.
|
||||
For example, suppose at t=0 with a Shortest Job First scheduler there are two processes (P1 P2) with 10 and 20 ms execution times.
|
||||
@@ -117,7 +117,7 @@ Here are measures of efficiency and their mathematical equations
|
||||
|
||||
\subsection{Convoy Effect}
|
||||
|
||||
The convoy effect is when a process takes up a lot of the CPU time, leaving all other processes with potentially smaller resource needs following like a Convoy Behind them.
|
||||
The convoy effect is when a process takes up a lot of the CPU time, leaving all other processes with potentially smaller resource needs following behind it like a convoy.
|
||||
|
||||
Suppose the CPU is currently assigned to a CPU intensive task and there is a set of I/O intensive processes that are in the ready queue.
|
||||
These processes require a tiny amount of CPU time but they are unable to proceed because they are waiting for the CPU-intensive task to be removed from the processor.
|
||||
@@ -127,11 +127,11 @@ For example, in the case of an FCFS scheduler, we must wait until the process is
|
||||
The I/O intensive processes can now finally satisfy their CPU needs, which they can do quickly because their CPU needs are small and the CPU is assigned back to the CPU-intensive process again.
|
||||
Thus the I/O performance of the whole system suffers through an indirect effect of starvation of CPU needs of all processes.
|
||||
|
||||
This effect is usually discussed in the context of FCFS scheduler; however, a Round Robin scheduler can also exhibit the Convoy Effect for long time-quanta.
|
||||
This effect is usually discussed in the context of FCFS scheduler; however, a Round Robin scheduler can also exhibit the convoy effect for long time-quanta.
|
||||
|
||||
\section{Scheduling Algorithms}
|
||||
|
||||
Unless otherwise stated
|
||||
Unless otherwise stated, the examples in this section use the following processes:
|
||||
|
||||
\begin{enumerate}
|
||||
\item Process 1: Runtime 1000ms
|
||||
@@ -258,7 +258,7 @@ Under SRTF (shortest remaining time first), which compares remaining rather than
|
||||
Processes are scheduled in the order of arrival.
|
||||
One advantage of FCFS is that the scheduling algorithm is simple.
|
||||
The ready queue is a FIFO (first in first out) queue.
|
||||
FCFS suffers from the Convoy effect.
|
||||
FCFS suffers from the convoy effect.
|
||||
Here P2 arrives, then P1 arrives, then P5, then P4, then P3.
|
||||
You can see the convoy effect for P5.
|
||||
|
||||
|
||||
@@ -34,7 +34,7 @@ If not possible, try to go through the engineering steps.
|
||||
You can't solve a problem that you don't fully understand.
|
||||
\item Determine whether you need to ``hack'' the system.
|
||||
A hack is defined generally as trying to use a system unintendedly.
|
||||
First, you should determine if your use is intended or unintended or somewhere in the middle -- get a decision from them.
|
||||
First, you should determine if your use is intended or unintended or somewhere in the middle -- get a decision from the system's owners.
|
||||
If you can't get that, make a reasonable judgement as to what the intended use is.
|
||||
\item Figure out a reasonable estimate of what the cost is to ``hacking'' the system.
|
||||
Get that reasonable estimate checked out with a few engineers so they can highlight things that you may have missed.
|
||||
@@ -48,7 +48,7 @@ This is often called a \keyword{policy vacuum}.
|
||||
This may seem like busy work and more on the ``business side'' than computer scientists are used to, but your career is at stake here.
|
||||
It is up to you as a computing professional to assess the risk and to decide whether to execute.
|
||||
Courts generally like sitting on precedent, but you can easily say that you aren't a legal scholar.
|
||||
In lieu, you must be able to say that you reacted as a ``reasonable'' engineer would react.
|
||||
Instead, you must be able to say that you reacted as a ``reasonable'' engineer would react.
|
||||
|
||||
\todo{Link to some case studies of real engineers having to decide}
|
||||
|
||||
@@ -162,6 +162,7 @@ int main() {
|
||||
\end{lstlisting}
|
||||
|
||||
\subsection{Out of order instructions \& Spectre}
|
||||
\label{sec:spectre}
|
||||
|
||||
Out of order execution is an amazing development that has been recently adopted by many hardware vendors (think 1990s) \todo{citation needed}.
|
||||
Processors now instead of executing a sequence of instructions (let's say assigning a variable and then another variable) execute instructions before the current one is done \cite[P. 45]{guide2011intel}.
|
||||
@@ -225,7 +226,7 @@ This could include important information such as passwords, payment information,
|
||||
The user gets matched with either the owner, the group, or `everyone else', and their access to the file is limited using these bits.
|
||||
Note that permissions work slightly differently on directories compared to files.
|
||||
\item Capabilities.
|
||||
In addition to permissions on files, each user has a certain set of permissions that they can do.
|
||||
In addition to permissions on files, a process can be given a certain set of capabilities: specific privileged actions that it is allowed to perform.
|
||||
For a full list, you can check capabilities(7).
|
||||
In short, allowing a capability allows a user to perform a set of actions.
|
||||
Some examples include controlling networking devices, creating special files, and peering into IPC or interprocess communication.
|
||||
@@ -322,7 +323,7 @@ As more and more of our systems are hacked over the web, it is important to unde
|
||||
\item Identity Verification.
|
||||
In TCP, there is no way to verify the identity of who the program is connecting to.
|
||||
There are no checks or federated databases in place.
|
||||
One just has to trust the DNS server gave a reasonable response which is almost always the incorrect answer.
|
||||
One just has to trust that the DNS server gave the correct address, and blindly trusting it is almost always the wrong decision.
|
||||
Apart from systems that have an approved white list or a ``secret'' connection protocol, there is little at the TCP level that one can do to stop this.
|
||||
\item Syn-Ack Sequence Number.
|
||||
This is a security improvement.
|
||||
|
||||
+4
-3
@@ -4,7 +4,7 @@
|
||||
|
||||
Signals are a convenient way to deliver low-priority information and for users to interact with their programs when other ways don't work (for example standard input being frozen).
|
||||
They allow a program to clean up or perform an action in the case of an event.
|
||||
Sometimes, a program can choose to ignore events which is supported.
|
||||
A program can also choose to ignore most signals, but \keyword{SIGKILL} and \keyword{SIGSTOP} can never be caught, blocked, or ignored.
|
||||
Crafting a program that uses signals well is tricky due to how signals are handled.
|
||||
As such, signals are usually for termination and clean up.
|
||||
Rarely are they supposed to be used in programming logic.
|
||||
@@ -43,7 +43,7 @@ The overall process for how a kernel sends a signal is below.
|
||||
This tells the kernel that when the process gets signal X that it should jump to function Y.
|
||||
\item A signal that is created is in a "generated" state.
|
||||
\item The time between when a signal is generated and the kernel can apply the mask rules is called the pending state.
|
||||
\item Then the kernel checks the process' signal mask.
|
||||
\item Then the kernel checks the process's signal mask.
|
||||
If the mask says all the threads in a process are blocking the signal, then the signal is currently blocked and nothing happens until a thread unblocks it.
|
||||
\item If a single thread can accept the signal, then the kernel executes the action in the disposition table.
|
||||
If the action is a default action, then no threads need to be paused.
|
||||
@@ -257,7 +257,8 @@ The \keyword{sig\_atomic\_t} type implies that all the bits of the variable can
|
||||
It is impossible to read a value that is composed of some new bit values and old bit values.
|
||||
|
||||
By specifying \keyword{pleaseStop} with the correct type \keyword{volatile\ sig\_atomic\_t}, we can write portable code where the main loop will be exited after the signal handler returns.
|
||||
The \keyword{sig\_atomic\_t} type can be as large as an \keyword{int} on most modern platforms but on embedded systems can be as small as a \keyword{char} and only able to represent (-127 to 127) values.
|
||||
The \keyword{sig\_atomic\_t} type can be as large as an \keyword{int} on most modern platforms but on embedded systems can be as small as a \keyword{char}.
|
||||
The C standard only guarantees that it can hold the values -127 to 127 if it is signed, or 0 to 255 if it is unsigned.
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
volatile sig_atomic_t pleaseStop;
|
||||
|
||||
@@ -227,7 +227,7 @@ Firstly, C Mutexes do not lock variables.
|
||||
A mutex is a simple data structure.
|
||||
It works with code, not data.
|
||||
If a mutex is locked, the other threads will continue.
|
||||
It's only when a thread attempts to lock a mutex that is already locked, will the thread have to wait.
|
||||
A thread only has to wait when it attempts to lock a mutex that is already locked.
|
||||
As soon as the original thread unlocks the mutex, the second (waiting) thread will acquire the lock and be able to continue.
|
||||
The following code creates a mutex that does effectively nothing.
|
||||
|
||||
@@ -397,8 +397,11 @@ When we wake up we try to grab the lock again.
|
||||
Once we successfully swap, we are in the critical section!
|
||||
We set the mutex's owner to the current thread for the unlock method and return successfully.
|
||||
|
||||
How does this guarantee mutual exclusion? When working with atomics we are unsure!
|
||||
But in this simple example, we can because the thread that can successfully expect the lock to be UNLOCKED (0) and swap it to a LOCKED (1) state is considered the winner.
|
||||
How does this guarantee mutual exclusion?
|
||||
Reasoning about atomics is often tricky, but this simple example is easy to check.
|
||||
The compare-and-swap is atomic, so only one thread can change the lock from UNLOCKED (0) to LOCKED (1).
|
||||
That thread's compare-and-swap succeeds and it enters the critical section.
|
||||
Every other thread sees LOCKED instead of the UNLOCKED value it expected, so its compare-and-swap fails and it keeps waiting.
|
||||
How do we implement unlock?
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
@@ -465,7 +468,7 @@ This becomes especially useful if you want to use a semaphore to implement a mut
|
||||
A mutex is a semaphore that always \keyword{waits} before it \keyword{posts}.
|
||||
Some textbooks will refer to a mutex as a binary semaphore.
|
||||
You do have to be careful to never add more than one to a semaphore or otherwise your mutex abstraction breaks.
|
||||
That is usually why a mutex is used to implement a semaphore and vice versa.
|
||||
With that care, a semaphore can stand in for a mutex, and later in this chapter we will go the other way and build a semaphore out of a mutex and a condition variable.
|
||||
|
||||
\begin{itemize}
|
||||
\item Initialize the semaphore with a count of one.
|
||||
@@ -526,7 +529,8 @@ Also, binary semaphores are different than mutexes because a mutex must be unloc
|
||||
|
||||
\subsubsection{Signal Safety}
|
||||
|
||||
Also, \keyword{sem\_post} is one of a handful of functions that can be correctly used inside a signal handler \keyword{pthread\_mutex\_unlock} is not.
|
||||
Also, \keyword{sem\_post} is one of a handful of functions that can be correctly used inside a signal handler.
|
||||
\keyword{pthread\_mutex\_unlock} is not.
|
||||
We can release a waiting thread that can now make all of the calls that we disallowed to call inside the signal handler itself e.g. \keyword{printf}.
|
||||
Here is some code that utilizes this:
|
||||
|
||||
@@ -1030,7 +1034,7 @@ There are three main desirable properties that we desire in a solution to the cr
|
||||
\begin{enumerate}
|
||||
\item Mutual Exclusion. The thread/process gets exclusive access.
|
||||
Others must wait until it exits the critical section.
|
||||
\item Bounded Wait. A thread/process cannot get superseded by another thread infinite amounts of time.
|
||||
\item Bounded Wait. A thread/process waiting to enter the critical section cannot be overtaken by other threads an unbounded number of times.
|
||||
\item Progress. If no thread/process is inside the critical section, the thread/process should be able to proceed without having to wait.
|
||||
\end{enumerate}
|
||||
|
||||
@@ -2262,7 +2266,7 @@ void* pop_elem(linked_list *ll, size_t index){
|
||||
\end{lstlisting}
|
||||
|
||||
\item How tight can you make the critical section?
|
||||
\item What is a producer consumer problem? How might the above be a producer consumer problem be used in the above section? How is a producer consumer problem related to a reader writer problem?
|
||||
\item What is a producer consumer problem? How might a producer consumer queue be used in the above section? How is a producer consumer problem related to a reader writer problem?
|
||||
\item What is a condition variable? Why is there an advantage to using one over a \keyword{while} loop?
|
||||
\item Why is this code dangerous?
|
||||
|
||||
@@ -2352,7 +2356,7 @@ void withdraw(int amount) {
|
||||
}
|
||||
\end{lstlisting}
|
||||
\item
|
||||
Sketch how to use a binary semaphore as a mutex. Remember in addition to mutual exclusion, a mutex can only ever be unlocked by the thread who called it.
|
||||
Sketch how to use a binary semaphore as a mutex. Remember in addition to mutual exclusion, a mutex can only ever be unlocked by the thread that locked it.
|
||||
\begin{lstlisting}[language=C]
|
||||
sem_t sem;
|
||||
|
||||
|
||||
+11
-11
@@ -145,7 +145,7 @@ The pthread library will automatically finish the process if no other threads ar
|
||||
\keyword{pthread\_exit(...)} is equivalent to returning from the thread's function; both finish the thread and also set the return value (void *pointer) for the thread.
|
||||
Calling \keyword{pthread\_exit} in the \keyword{main} thread is a common way for simple programs to ensure that all threads finish.
|
||||
For example, in the following program, the \keyword{myfunc} threads will probably not have time to get started.
|
||||
On the other hand \keyword{exit()} exits the entire process and sets the process' exit value.
|
||||
On the other hand \keyword{exit()} exits the entire process and sets the process's exit value.
|
||||
This is equivalent to \keyword{return ();} in the main method.
|
||||
All threads inside the process are stopped.
|
||||
Note the \keyword{pthread\_exit} version creates thread zombies; however, this is not a long-running process, so we don't care.
|
||||
@@ -200,9 +200,9 @@ Here is a non-complete list:
|
||||
|
||||
\section{Race Conditions}
|
||||
|
||||
Race conditions are whenever the outcome of a program is determined by its sequence of events determined by the processor.
|
||||
A race condition is when the outcome of a program depends on the order in which its threads happen to run, which is decided by the scheduler.
|
||||
This means that the execution of the code is non-deterministic.
|
||||
Meaning that the same program can run multiple times and depending on how the kernel schedules the threads could produce inaccurate results.
|
||||
The same program can run multiple times and, depending on how the kernel schedules the threads, produce different (and incorrect) results.
|
||||
The following is the canonical race condition.
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
@@ -227,12 +227,12 @@ int main() {
|
||||
\end{lstlisting}
|
||||
|
||||
Breaking down the assembly there are many different accesses of the code.
|
||||
We will assume that data is stored in the \keyword{eax} register.
|
||||
The code to increment is the following with no optimization (assume int\_ptr contains eax).
|
||||
For simplicity, assume the shared integer lives in memory at \keyword{[rbp-4]}.
|
||||
Without optimization, the update compiles to something like the following: load the value from memory into the \keyword{eax} register, double it, and store it back.
|
||||
|
||||
\begin{lstlisting}[language={[x86masm]Assembler}]
|
||||
mov eax, DWORD PTR [rbp-4] ;Loads int_ptr
|
||||
add eax, eax ;Does the addition
|
||||
mov eax, DWORD PTR [rbp-4] ;Loads the value from memory
|
||||
add eax, eax ;Doubles it
|
||||
mov DWORD PTR [rbp-4], eax ;Stores it back
|
||||
\end{lstlisting}
|
||||
|
||||
@@ -301,11 +301,11 @@ The above code suffers from a \keyword{race condition} - the value of i is chang
|
||||
The new threads start later; in the example output, the last thread starts after the loop has finished.
|
||||
To overcome this race-condition, we will give each thread a pointer to its own data area.
|
||||
For example, for each thread we may want to store the id, a starting value and an output value.
|
||||
We will instead treat i as a pointer and cast it by value.
|
||||
Alternatively, for a single small integer, we can cast the value of \keyword{i} to a \keyword{void*} and pass that value directly.
|
||||
|
||||
\begin{lstlisting}[language=C]
|
||||
void* myfunc(void* ptr) {
|
||||
int data = ((int) ptr);
|
||||
int data = (int) (intptr_t) ptr;
|
||||
printf("%d ", data);
|
||||
return NULL;
|
||||
}
|
||||
@@ -315,7 +315,7 @@ int main() {
|
||||
int i;
|
||||
pthread_t tid;
|
||||
for(i =0; i < 10; i++) {
|
||||
pthread_create(&tid, NULL, myfunc, (void *)i);
|
||||
pthread_create(&tid, NULL, myfunc, (void *) (intptr_t) i);
|
||||
}
|
||||
pthread_exit(NULL);
|
||||
}
|
||||
@@ -548,7 +548,7 @@ Guiding questions:
|
||||
\item What is the first argument to pthread create?
|
||||
\item What is the start routine in pthread create? How about arg?
|
||||
\item Why might pthread create fail?
|
||||
\item What are a few things that threads share in a process? What are a few things that threads have different?
|
||||
\item What are a few things that threads share in a process? What are a few things that differ between threads?
|
||||
\item How can a thread uniquely identify itself?
|
||||
\item What are some examples of non thread safe library functions? Why might they not be thread safe?
|
||||
\item How can a program stop a thread?
|
||||
|
||||
Reference in New Issue
Block a user