EmC Language
EmC is a statically-typed, heap-less, C-style scripting language for memory-constrained 32-bit embedded systems. Source files (.emc) compile to a compact bytecode binary (.emcbin) that runs on a stack-based virtual machine with no external dependencies and no runtime heap allocation.
The syntax is close to C. Most C code that avoids pointers, the heap, and the standard library will look familiar.
1. Program Structure
There is no main function. Top-level statements execute in source order when the script runs.
System:: is the namespace for functions provided by the host application, called native functions. A host can also group natives under namespaces of its own. See Native functions.
Functions and classes may be declared at the top level. They can be referenced before their declaration appears in the file (see Forward references).
2. Comments
EmC supports C-style single line and block comments.
3. Types
EmC is a statically typed language. All data types are known at compile time, making execution faster and safer.
| Keyword | Alias | Meaning | Width |
|---|---|---|---|
void | no value (return type) | - | |
bool | true / false | 1 byte | |
char | s8 | signed 8-bit integer | 1 byte |
byte | u8 | unsigned 8-bit integer | 1 byte |
short | s16 | signed 16-bit integer | 2 byte |
ushort | u16 | unsigned 16-bit integer | 2 byte |
int | s32 | signed 32-bit integer | 4 byte |
uint | u32 | unsigned 32-bit integer | 4 byte |
float | f32 | 32-bit IEEE-754 | 4 byte |
string | immutable text constant | ref |
The alias column is just another way to declare the same keyword: s8 and char compile to exactly the same type, so they mix freely and error messages always report the char/byte/… name.
Aliases are provided for convenience as they are a little more descriptive than the classical type names.
char and byte are numeric types, not a distinct character type. Individual character literals are not supported, so an integer must be used.
string is a reference to a compile-time constant. String variables and string native parameters can only ever hold a literal or another string constant. For mutable text, use a byte array (see Strings).
4. Literals
null, NULL, and nil are all accepted and evaluate to integer 0.
An integer literal without a suffix is an int. Add U to make it a uint, as in C. This matters in two places: a plain literal assigned to a uint variable warns about an implicit cast, and an expression is only compiled with unsigned arithmetic and comparisons when one of its operands is unsigned, so x > 0U and x / 2U behave correctly for a uint x above 2147483647. A literal larger than 4294967295 is a compile error. An unsuffixed literal above 2147483647 keeps its bit pattern and wraps negative as an int, with a warning unless it is being assigned to a uint (so uint mask = 0xFFFFFFFF; is fine) or negated.
String literals use double quotes. There are no escape sequences: "\n" is a backslash followed by n, not a newline.
5. Variables
All variables are zero-initialized if no initializer is given.
Scope
- Variables declared at the top level are globals.
- Variables declared inside a function or block are locals.
- Blocks (
{ ... }) introduce a new scope. A local is visible from its declaration to the end of its enclosing block.
const
const marks a variable read-only after its initializer runs. Writing to it later is a compile error.
A const array must be initialized where it is declared, and its elements can’t be written after that. Assignment, compound assignment and ++/-- on an element are all compile errors. This works well for lookup tables:
A const array can only be passed to a const array parameter (see “Array Parameters”).
Naming Style
EmC enforces no naming convention. One common pattern is UpperCamelCase for functions and global variables and lowerCamelCase for locals and parameters, which makes a name’s scope visible at a glance and matches how the native functions are named. The compiler accepts any style, and the choice is entirely yours.
6. Operators
All basic mathematical and bitwise operations are supported.
Arithmetic
Important
Integer division and modulus by zero halt the VM with a “Division By Zero” error. Float division by zero also halts.
Bitwise
Comparison
A comparison always produces a bool.
Logical
Assignment
There is no %=, <<=, or >>=.
Increment / Decrement
Both prefix and postfix forms work on a variable, a class field or an array element:
A postfix ++ or -- has to be the whole expression, as in i++; or int old = counts[b]++;. Using it inside a larger expression, such as a + i++, is a compile error. The index is evaluated once, so counts[Next()]++; calls Next() once.
Narrow types wrap at their own width:
Ternary
Behaves like a single line if/else statement.
Precedence
From tightest to loosest binding:
.[]()(member, index, call)!~++--(unary)*/%+-&|^<<>><><=>===!=&&||?:(ternary)=+=-=*=/=&=|=^=(assignment)
Note that the bitwise and shift operators sit at the same level as + and -. Parenthesize when mixing them with arithmetic.
7. Type Conversions
Numeric types convert implicitly across signed, unsigned, and float categories as needed by an assignment, argument, or operator. There is no explicit cast syntax.
A function call’s return value converts the same way as any other value, so float f = count(); converts an int result to float and int n = 2 * ratio(); converts a float result to int.
The compiler warns when a float is converted to an integer type, because the fractional part is lost. This covers declarations, assignments, arguments, return values and array indexes. Other conversions, such as signed to unsigned, don’t warn by default.
Mixing int and float in arithmetic gives a float, as in C. The value keeps its fraction until it is stored, passed or returned, so int n = i * f * 4; only converts to int once, at the end. When the target is a float, the whole expression is worked out in float, including integer division. So float h = i / 2; with i = 1 gives 0.5, not 0.
Bitwise operators and switch need integer values, so a mixed expression like (i + f) & 3 is a compile error.
Storing a value in a char, byte, short or ushort wraps it to that type’s range, as in C. That happens at every store: a declaration, an assignment, a compound assignment, ++/--, an argument, and a return value. Arithmetic in between is done in int, so a result is only cut down when it’s stored. This makes decoding a signed 16-bit value from two bytes work as expected:
There is no integer overflow or wraparound detection. Arithmetic that exceeds a type’s range wraps silently.
8. Control Flow
if / else if / else
while
A while loop will continue as long as the condition is true.
There is no do/while.
for
A for loop has an initializer, condition, and update expression. The loop will continue as long as the condition is true.
All three clauses are optional. Any of them may be left empty, and for (;;) is an infinite loop.
break & continue
continuewill skip the rest of the loop body and continue to the next loop iteration.breakexits the loop immediately.
break and continue work in while and for loops.
switch
The controlling expression is an integer (float is rejected). case labels are integer literals and must be unique. default is optional. A case without a break falls through to the next case, as in C.
A case with no body falls straight into the next one, which is how several values share a single handler:
A case that has a body but no break runs its own body and then continues into the next case:
Info
switch statements are much more efficient than long if/else chains. They use a jump table to very quickly jump to the correct case label, rather than checking every entry for a match. The tradeoff is they produce more compiled binary size for large ranges.
A switch statement has to produce a jump table entry for every value between its lowest and highest case label value. This means that for large value ranges with a low number of case labels, the compiled output will be huge compared to if/else. The number of case labels has no effect on performance.
The compiler will output a warning if the number of case labels is less than half the value range.
9. Functions
- Scalar parameters are passed by value.
- A non-
voidfunction mustreturna value. - Calling with the wrong number of arguments is a compile error.
- Recursion is supported.
- A bare
return;in top-level code ends the script early. Destructors still run for every class instance declared before it. Returning a value from top-level code is a compile error.
Forward References
A top-level function or class may be used before it is declared in the file. This allows mutual recursion:
Array Parameters
An array parameter is written T name[] or T *name (equivalent). The array’s size is not part of the parameter type, so pass the length as a separate argument. A variable index used inside the function is still checked at runtime against the caller’s actual array, the same as any other array (see “Bounds checking” in the Arrays section).
Passing a non-array where an array parameter is expected, or an array of the wrong element type, is a compile error. So is giving the parameter a size (int values[4]).
Inside the function the parameter is used exactly like an array: index it, or pass it on by name to another array parameter, script or native. It cannot be used as a value on its own.
Declare an array parameter const when the function only reads it. A const parameter can’t be written inside the function, and it accepts both const and ordinary arrays. A const array can only be passed to a const parameter, since an ordinary parameter could write to it. The same applies when a function passes its own const parameter on.
Class parameters
See Classes.
10. Arrays
- Array size is fixed at compile time.
- If an initializer list is present, its length must match the declared size exactly.
int a[3] = {1, 2};is a compile error. - Without an initializer, every element is zero.
Indexing and assignment
Packed storage
char/byte arrays pack 4 elements per 4-byte slot; short/ushort arrays pack 2 per slot. This is transparent to the script; index them normally.
Bounds checking
A literal out-of-range index is a compile error, including a negative one:
A variable or computed index is checked at runtime instead. An out-of-range access halts the VM with an “Array Index Out Of Bounds” error rather than reading or writing whatever happens to sit next to the array:
This covers array parameters too: a function indexing a T name[] parameter with a variable index is checked against the size of whatever array the caller actually passed in, even through several levels of forwarding. It also covers an array field declared inside a class, however it’s reached: directly, through a class-typed parameter, or through a composed/embedded instance. It does not cover what a native function does with an array you pass it. See Native functions.
Bare array references
An array name used without an index has no value. It cannot be assigned to a scalar, returned as a scalar, or passed as a scalar argument. Pass it only to an array parameter (with a length) or index it.
11. Strings
String literals are immutable compile-time constants. Use them directly with the print natives:
For text you need to build or modify at runtime, use a byte buffer and the string natives. Every string native takes an explicit capacity, C snprintf-style. Nothing grows a buffer for you.
Notes and limits:
StrCopy/StrAppendtake astringconstant as the source, not anotherbyte[]buffer.- A freshly declared buffer is zero-filled, so
StrLengthon an untouched buffer is 0. - Content that does not fit the capacity is truncated and null-terminated.
12. Classes
Classes are supported for advanced data structures.
Fields
Declared in the class body. Each instance gets its own copy.
Methods
- Inside a method,
this.fieldand a barefieldname both refer to the current instance’s field. - A method can call a sibling method on the same instance with
this.method().
Constructors
ClassName(params) { ... }. A class has at most one constructor. It runs when an instance is declared with an argument list:
Fields are always zero-initialized first, before the constructor body runs. A class with no constructor is declared without parentheses:
Declaring an instance of a class that has a constructor without an argument list compiles, but produces a warning.
Destructors
~ClassName() { ... }. Runs automatically when the instance goes out of scope:
- A local instance is destroyed at the end of its enclosing block.
- Multiple instances in the same scope are destroyed in reverse declaration order (LIFO).
- Global instances are destroyed once, at script end, in reverse declaration order.
Passing instances
Pass an instance to a function or method by reference with ClassName *param. The callee can call methods on it and read or write its fields, including compound assignment.
Passing an instance of the wrong class, or a non-instance, is a compile error. A bare instance name (no ., no method call, not passed to a class parameter) has no value and cannot be used as one.
Class-typed fields (composition)
A field may itself be a class instance. The embedded instance is laid out inline in its owner and reached with a chain of .:
- Nesting is unlimited:
a.b.c.xresolves as long as each step names a class-typed field. - Compound assignment works through the chain:
o.inner.value += 50;. - An embedded instance can be passed by reference like any other instance:
take(o.inner);wheretaketakesInner *i. - When an instance is created, each embedded field is zero-initialized and its field-default initializers run, outermost first.
- Destructors run automatically and in order: the owner’s destructor body first, then each embedded field’s destructor in reverse declaration order.
Limits:
- No member-initializer syntax. You cannot pass constructor arguments to an embedded field. Embedding a class whose constructor takes arguments is a compile error. A class with no constructor (or a parameterless one) is fine.
- No cycles. A class cannot contain itself, directly or indirectly (
class A { A a; }, orAholding aBthat holds anA). This is a compile error.
Not supported
- Inheritance. Every class is standalone. There is no subclassing.
- Class-typed return values. A function cannot return a class instance.
13. Namespaces
Namespaces are optional. They group top-level declarations under a name so they can be kept tidy and referred to explicitly. They have no runtime cost or effect: a namespaced global is still a plain global, and a namespaced function is still an ordinary function.
Refer to a member from outside with the :: scope operator:
- Unqualified access inside the block. Within
namespace Geometry { }, other members are visible without the prefix (area()can callgridSizeandPointdirectly). Names that don’t resolve inside the namespace fall back to the global scope. - Reopening. The same namespace name may be opened more than once, and the contents are merged.
- Forward references work across the whole file, exactly as they do at the top level.
- No nesting. A
namespacecannot be declared inside anothernamespace. - Native namespaces are reserved. A script cannot declare a namespace that the host’s natives use, and
Systemis always off limits. See Native functions. Mathis reserved for the built-in math functions. See Math functions.
14. Preprocessor
Runs on the token stream before parsing. Two directives are supported.
#include
- Path is resolved relative to the including file first. If it is not found there, any include directories set up by the host are searched in order. Absolute paths are used as is.
- Each file is included at most once, so diamond includes are safe.
- A circular include is a compile error, not a hang.
- A missing file is a compile error.
- Errors inside an included file are reported against that file’s own line numbers.
#define
Object-like macros only.
- A name is a macro only from its
#defineonward. - A macro body may reference an earlier macro, which is re-scanned and expanded.
- Self-referential and mutually-referential macros expand once and stop.
- An empty replacement is allowed and vanishes at the use site.
- Function-like macros (
#define SQ(x) ((x)*(x))) are a compile error. - There are no conditional directives (
#ifdef,#if,#endif,#undef). - A macro defined in an including file is visible inside included files.
15. Native functions
Native functions are provided by the host application. They give a script access to the device it runs on, for example printing, timing and communications. Which natives are available depends on the host. The reference set below is a common starting point.
Calling natives
Every native must be called through its namespace:
A bare Yield(10) is a compile error.
Most natives are in System. A host can group others under namespaces of its own, such as CAN above.
- A namespace used by any native is reserved, so a script cannot declare one with the same name.
Systemis always reserved. - Only the namespace is reserved, not the names inside it. With
CAN::Readavailable, a script is still free to declare its ownReadvariable or function, or aData::Readof its own.
If the script calls a native the host doesn’t provide, it still compiles, but halts with a “Native Function Not Resolved” error when the call runs.
Array arguments
Pass an array to a native by bare name, followed by its length:
The native trusts the length you give it. Its access to the array is not bounds checked, so never pass a length larger than the array.
A native that only reads an array takes it as a const parameter, and accepts both const and ordinary arrays. Passing a const array to a native that isn’t marked const is a compile error.
Callback parameters
Some natives take a script function as an argument, so the host can call back into the script later, for example when a CAN frame arrives. Pass the function’s bare name, with no parentheses:
A callback always returns void, and its parameters must match what the native expects. The compiler checks the parameter count, each parameter’s type, and whether it is an array, and reports a mismatch at the call site. Some natives accept any void function. A wrong parameter count is then only caught when the host calls it, which halts the script with a “Call Arg Count Error”.
An array parameter in a callback is only valid for the duration of the call. Index it or pass it on to another array parameter as usual, and copy anything you want to keep into a script array before returning.
Reference set
These natives are all in the System namespace, so Print is called as System::Print("hi").
| Signature | Purpose |
|---|---|
void SetError(int code) | Signal a recoverable error code to the host. |
void Print(string str) | Write a string, no newline. |
void PrintLine(string str) | Write a string and a newline. |
void PrintInt(int i) | Write an integer and a newline. |
void PrintFloat(float f) | Write a float and a newline. |
void PrintFormat(string str, float f) | Write a format string with one float (PrintFormat("v: %f", 3.14)). |
int StrLength(const byte buf[], int capacity) | Length up to the null terminator or capacity. |
void StrCopy(byte dest[], int destCapacity, string src) | Copy a string constant into a buffer, truncating to fit. |
void StrAppend(byte dest[], int destCapacity, string src) | Append a string constant onto a buffer’s content. |
void IntToStr(byte dest[], int destCapacity, int value) | Format an integer as decimal text into a buffer. |
bool StrEquals(const byte a[], int capA, const byte b[], int capB) | Compare two buffers’ null-terminated contents. |
void PrintBuffer(const byte buf[], int capacity) | Write a buffer’s null-terminated content. |
uint NowMs() | Host uptime in milliseconds. Wraps on overflow - compare with unsigned subtraction. |
uint NowUs() | Host uptime in microseconds. Wraps on overflow, typically much sooner than NowMs. |
void Yield(uint t) | Pause for t milliseconds. |
uint YieldUntil(uint lastTime, uint delay) | Fixed-period pause: sleeps until lastTime + delay, returns the new lastTime to pass back in next iteration. |
SetError
SetError(int) records an error code without halting the VM. The host application can read it after the script finishes. The last call wins. The default is 0.
16. Math functions
A set of math functions is built into the language under the Math namespace. They are part of the VM itself, so they are always available, whatever natives the host provides.
| Function | Result |
|---|---|
Math::Sqrt(x) | Square root. |
Math::Pow(base, exp) | base raised to exp. |
Math::Sin(x), Math::Cos(x), Math::Tan(x) | Trigonometric functions. x is in radians. |
Math::Asin(x), Math::Acos(x), Math::Atan(x) | Inverse trigonometric functions, in radians. |
Math::Atan2(y, x) | Angle of the point (x, y) in radians, in the range -pi to pi. |
Math::Exp(x) | e raised to x. |
Math::Log(x) | Natural logarithm. |
Math::Log2(x), Math::Log10(x) | Base 2 and base 10 logarithms. |
Math::Floor(x), Math::Ceil(x) | Round down / up to a whole number. |
Math::Round(x) | Round to the nearest whole number, halves away from zero. |
Math::Fmod(x, y) | Floating point remainder of x / y. |
Math::Abs(x) | Absolute value. |
Math::Min(a, b), Math::Max(a, b) | Smaller / larger of the two. |
All functions take and return float. An int argument is converted to float first, the same as passing it to a float parameter.
Abs, Min and Max are the exception: when every argument is an integer type they work in integers and return int, so int m = Math::Max(3, 7); is exact with no float round trip. If any argument is a float, including a mixed expression like i + 0.5, the whole call is done in float.
A domain error such as Math::Sqrt(-1.0) or Math::Log(0.0) produces NaN or infinity, as in C. It does not halt the VM.
17. Runtime Model and Limits
- No heap. Globals, locals, and call frames all live in one buffer supplied by the host. The VM never allocates at runtime.
- No whole-script size ceiling. Function entry points are 32-bit offsets, so a compiled binary can be as large as the host is willing to load. Two 16-bit limits do apply: a single
if,else, loop body, orswitchcannot span more than 64 KB of bytecode (the compiler reports “Too much code to jump over”), and the VM’s data area (globals plus stack) is capped at 65,535 slots, which is 256 KB. - Fixed stack. The host sets the stack size. Deep recursion or large local arrays can exhaust it (“Stack Overflow”).
- Bounds checking. Stack overflow/underflow and out-of-range pointer dereferences are caught and halt the VM rather than corrupting memory. The host can disable this for targets that cannot afford the checks. A local, global, or array-parameter index that runs off the end of its array is also caught (“Array Index Out Of Bounds”). See “Bounds checking” under Arrays.
- Division by zero halts the VM (“Division By Zero”) for integer and float operands.
- No overflow detection. Arithmetic wraps silently.
- No exceptions. There is no
try/catch/throw. A genuine runtime error halts the VM. UseSetErrorfor recoverable conditions. - No string escape sequences.
- Language version. A compiled binary records the language version it was built for. A VM won’t load a binary from a newer version, for example a 0.3 script on a 0.2 VM, since it may use instructions that VM doesn’t have. Binaries from older versions still load. Recompile a script with the matching compiler, or update the VM.
Runtime errors
Every run ends with a status. “End” means the script finished normally. Anything else is an error that stopped it. The common ones are:
| Error | Cause |
|---|---|
| Division By Zero | An integer or float division, or a modulus, by zero. |
| Array Index Out Of Bounds | A variable array index was outside the array. |
| Stack Overflow | The script ran out of stack, usually from deep recursion or large local arrays. |
| Native Function Not Resolved | The script called a native the host doesn’t provide. |
| Call Arg Count Error | The host called a callback with the wrong number of arguments. |
| Stack Underflow, Pointer Out Of Bounds, Unknown Instruction | A damaged binary, or a problem in the VM or host rather than in the script. |
A binary can also be refused when it’s loaded, before any of it runs: “Unsupported Version” for a binary compiled for a newer language version, and “Invalid Script” for one that is malformed or fails its checksum.
18. Complete Example
Output: