Operating Systems 2026F Lecture 5

From Soma-notes
Revision as of 21:10, 25 September 2026 by Soma (talk | contribs) (Created page with "==Video== Video from the lectures given on September 24th and 25th, 2026 are now available: * [https://homeostasis.scs.carleton.ca/~soma/os-2026f/lectures/comp3000-2026f-lec05a-20260924.mp4 Lecture 5A] * [https://homeostasis.scs.carleton.ca/~soma/os-2026f/lectures/comp3000-2026f-lec05b-20260925.mp4 Lecture 5B] ==Notes== ===Lecture 5A=== <pre> Lecture 5a ---------- The memory map of a process - every process has its own view of memory, cannot see the memory of other...")
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)

Video

Video from the lectures given on September 24th and 25th, 2026 are now available:

Notes

Lecture 5A

Lecture 5a
----------

The memory map of a process
 - every process has its own view of memory, cannot see the memory of other processes
 - what goes in that address space?

Note that, almost always, the entire address space is NOT VALID MEMORY
 - if you access most of it, you'll get an error (segmentation error generally)

So what is a segment?

Well first, consider the address
 - 64 bit pointer, so 2^64 possible addresses

We can't have this much RAM, so most of these addresses are invalid.
 - but which ones are valid?
 - and on what basis?


First, we should talk about the stack and the heap

The heap is just memory that is allocated in chunks, with each chunk pointed to
by a pointer
 - when you call malloc, you're getting memory from the heap
 - when you create a new object in most languages other than C, you're really
   getting memory from the heap (mostly)

How to manage the heap is a complex problem, solved by memory allocators, garbage collection, etc (beyond the scope of this class)

The stack is memory that is managed in a very simple way: as a stack
 - LIFO (last in, first out), e.g., a stack of plates

A stack is great because...
 - no memory leaks!
 - allocation and de-allocation is trivial

But the discipline of a stack only makes sense in certain contexts

Fortunately, we've built our programming languages around those contexts:
 FUNCTIONS

When a program enters a function, it allocates memory for the function
When the function exits, that function's storage needs to be de-allocated

So we need somewhere to put the stack and somewhere to put the heap
 - could separate, but normally are allocated together

Stack grows from himem (of data segment) down
heap grows up from the bottom of the data segment

if you put some invalid memory in the middle, you can know when they would potentially overlap (because access to it will generate an error)

a segmentation violation is just an access to memory that hasn't been allocated

So what's in memory?

<top>
  command line args, environment vars

  Data
 ------ "break"
  Code

<bottom>


So the command line arguments and environment variables passed to a program
are at the top of the process's address space. The kernel put them there when
the program was launched.


fork duplicates the current process
execve replaces the current program with a new one
 - neither act like normal "functions", because they are system calls

Windows has 

Remember in C we make system calls by calling functions
 - the functions that make system calls are using compiler-specific ways
   of making the special CPU instructions for system calls



fork <- to make a process
execve <- to load a program into a process
exit <- to terminate a process
wait <- wait for a child to finish


If the parent doesn't wait for its child, and the child terminates,
the child sticks around in a zombie state
 - it is already dead but still around, cannot be further killed

zombies stick around as long as their parent exists
zombies are "reaped" once the parent dies
 - init or similar will call wait


How do we run programs?

Run fork
 - duplicates current process, creating a child process

in the parent
 - wait for the child to finish, or
 - go do other work and call wait will notified
   of child termination

in the child
 - set up things for new program
   - standard in, out, error, other open files
     - close any files that SHOULD NOT be accessible
   - command line arguments
   - environment variables
 - execve new program

In practice, you won't actually see the fork system call nowadays,
at least on Linux systems. Instead, you'll see clone


Containers - have you heard of them?

So containers are all about managing software and dependencies
 - very poor security isolation properties

Turns out when you create a new process, you can change its view of the world
 - e.g., give it a different view of the filesystem

namespaces is how you map the following for a process and its children

 /newsystem/bin => /bin
 uid 1000 => uid 0

Lecture 5B

Lecture 5b
----------

Planning to move midterm to Oct 20, 21
 - so review will be Oct 15, 16
 - assignment 2 will be due on the 15th 2:30 PM
 - a1 will be put off a bit will finalize when posted (you'll have at least a week)
 - A1 will be based on T1 and T2

For today, we're talking memory

With modern OSs, each running process gets its own address space
 - so 64 bits in size
 - most of that address space is not allocated, so cannot
   be used or will generate an error
 - typically the error is a "segmentation violation"
   - so what is a segment?

A segment is a unit of memory that has a specific "purpose"
 - contiguous addresses

What segments are there? You'll at least have
 - code (text), read only
 - data, read write
   - global variables
   - dynamically-allocated storage

The data segment has to store two logically distinct types of dynamically allocated memory: the stack and the heap

The heap is used to store dynamically allocated objects/data structures that have indeterminate lifetimes
 - could be around for a few milliseconds
 - could exist for most of the run of the program (but aren't known at compile time)
 - malloc (C), new (C++), most everything in JavaScript & Python

Now this might sound complicated because IT IS
 - entire area of memory management techniques
 - automatic memory management is called "garbage collection"
   - not only but most general approach

The heap is complex, hard to do right, so traditionally UNIX programs have avoided using the heap as much as possible and instead used the stack

So what's the stack?
 - data structure that respects LIFO discipline (last in, first out)

A process normally has one stack that is used for functions
 - local variables
 - arguments
 - return values

heap grows up from bottom of data segment
stack grows down from the top of the data segment

normally some invalid memory in the middle so you know when the two run into each other

Where do the environment variables and command line variables come from?

Those are both arguments of the execve system call


Key system calls for running a program in a new process are:
 - fork
 - execve
 - wait

fork: duplicates the current process
 - new process is called the child, old the parent
 - except for the PID/PPID and the return value of fork(),
   code, data, and other state are IDENTICAL

child calls execve
 - sets up argv, env, open files (standard in/out/error +),
   signal handlers (will discuss later)
 - then loads new program binary with execve call, which REPLACES
   the code in the process
 - if execve returns, it failed

parent calls wait
 - either literally waiting for child process to terminate, or
 - in response to a signal that the child has terminated

Why does the parent care about the child process terminating?
 - because it has a return status

processes terminate with the exit system call, and exit takes an integer argument that is past via wait to the parent

Modern Linux systems tend not to use the fork system call
 - fork() actually calls clone

If you don't use setarch -R to run 3000memview, addresses will be randomized. This is ASLR in action: address space layout randomization