diff options
| -rw-r--r-- | Learning_notes/Memristor & Memcapacitor in Neural networks.md | 12 | ||||
| -rw-r--r-- | Learning_notes/NeuroBench/algorithm track.md | 25 | ||||
| -rw-r--r-- | Learning_notes/NeuroBench/neuroBench_notes.md | 8 | ||||
| -rw-r--r-- | Learning_notes/NeuroBench/system track.md | 15 | ||||
| -rw-r--r-- | README.md | 2 | ||||
| -rw-r--r-- | custom_benchmark/README.md | 5 | ||||
| -rw-r--r-- | custom_benchmark/memTorch_SNN.py | 28 | ||||
| -rw-r--r-- | custom_benchmark/memristor_benchmark.cpp | 2 | ||||
| -rw-r--r-- | neurobench_testing/README.md | 1 | ||||
| -rw-r--r-- | neurobench_testing/custom_memristor_model.py | 18 | ||||
| -rw-r--r-- | neurobench_testing/kaevin_memristor.py | 134 | ||||
| -rw-r--r-- | neurobench_testing/memTorch_testing.py | 89 | ||||
| -rw-r--r-- | neurobench_testing/memTorch_testing_custom_memristor.py | 160 | ||||
| -rwxr-xr-x | quickstart | bin | 0 -> 138224 bytes | |||
| -rw-r--r-- | references/kaevin_memristor.py | 4 | ||||
| m--------- | spires | 0 | ||||
| -rw-r--r-- | spires_tutorial.c | 79 |
17 files changed, 306 insertions, 276 deletions
diff --git a/Learning_notes/Memristor & Memcapacitor in Neural networks.md b/Learning_notes/Memristor & Memcapacitor in Neural networks.md deleted file mode 100644 index bd1cc20..0000000 --- a/Learning_notes/Memristor & Memcapacitor in Neural networks.md +++ /dev/null @@ -1,12 +0,0 @@ -Just some quick notes on memory components in crossbar arrays for Neural Networks -## Von Neumann Bottleneck -The architecture strictly seperates the processing unit from the memory unit. Everytime a calculation is done in a neural network the processor has to fetch the "weights" from memory, do the calculation and then send it back. Modern neural networks can require billions or trillions of weights, so up to 90% of the energy and time spent are just moving data back and forth. memristor & memcapacanges depending on the history of the voltage and current that have previously paitive circuits eliminate this by doing in memory computing (IMC). - -## memristors -A memristor's resistance changes depending on the history of the voltage and current that have previously passed through it. So it has non volatile memory. By applying specific voltage pulses you can tune the resistance to a high resistance state or a low resistance state, essentially giving it on and off states. - -Neural networks use Vector Matrix multiplication (VMM). Where a set of inputs is multiplied by a matrix of weights to determine how strongly a set of neurons should fire. Instead of having to calculate "line by line" in CMOS NN's memristor networks arrange thousands of memristors into a crossbar array (basically a grid) and can solve VMM instantly and in parallel. It does this using ohm's law and kirchoff's current law. When you apply an input voltage, it passes through a memristor that has a certain conductance, and the resulting electrical current is the multiplication of the input (V x G) then when all the currents meet in a single column they add together and that is the weighted sum. - -## memcapacitors -Have practically zero static power dissipation. These are great for reservoir computing, a type of Recurrent Neural Network (RNN) used to process time-series data, things like speech recognition, heartbeat monitoring, or epilepsy detection. Also you can read the resovoir state without needing an extra read pulse so they are faster and more efficient. -Because a memcapacitor dynamically changes how it accumulates charge based on its past stimulation it perfectly replicates short-term and long-term plasticity just like cell membranes for memory. We can arrange these in a crossbar array but instead of measuring the output current we have to measure the displacement charge. So the NN can still perform VMM but with a much smaller thermal footprint. It does have some difficulties though, Reading the displacement charge is technically more difficult than just an output current, and actually making memcapacitors is more difficult than resistive RAM (RRAM).
\ No newline at end of file diff --git a/Learning_notes/NeuroBench/algorithm track.md b/Learning_notes/NeuroBench/algorithm track.md deleted file mode 100644 index aaaedc8..0000000 --- a/Learning_notes/NeuroBench/algorithm track.md +++ /dev/null @@ -1,25 +0,0 @@ -aims to evaluate algorithms in a system independant manner. This means that the implementation platform can be ill matched to the algorithm benchmark it executes. Which allows for the algorithm complexity and expected performance to be examined in a theoretical manner, resulting in quicker prototyping and and functional analysis. - -## metrics -**footprint:** A measure of the memory footprint in bytes. This metric summarizes synaptic weight count, weight precision, trainable neuron paramaters, data buffers, etc... Zero weight are included because they are distinguished in the connection sparsity metric. - -**Connection sparsity:** is the the number of zero weights divided by the total number of weights, accumulated over all layers. 0 refers to no sparsity (fully connected) and 1 means no sparsity. - -**Activation sparsity:** The average sparsity of neuron activations over all neurons in all model layers, for all timesteps of all tested samples. again 0 refers to no sparsity (all neurons are activated) and 1 refers to the case where all neurons have a zero output. - -**Synaptic Operations:** Average number of synaptic operation per model execution, based on neuron activations and the associated fanout synapses. The metric can be subdivided into dense, effective multiply-accumulate, and effective accumulate synaptic operations (dense, Eff_MACs, Eff_ACs). Dense accounts for all zero and non-zero neuron activations and synaptic connections, and reflects the number of operations necessary on hardware that does not support sparsity. Eff_MACs and Eff_ACs only count effective synaptic operations by disregarding zero activations and zero connections, thus reflecting operation cost on sparsity-aware hardware. Synaptic operations with non-binary activation are considered multiply accumulates(Eff_MACs), while those with binary activation are considered accumulates (ACs). - -Footprint & connection sparsity are static metrics, so they can be determined from the model alone. activation sparsity, synaptic operations, and correctness are classified as workload metrics, so they depend on the execution of the model based on the benchmark data. -## benchmarks -**FSCIL:** Few-Shot Class Incremental Learning (FSCIL) is learning new tasks from a small amount of experiences while retaining knowledge of prior tasks is a key characteristic of biological intelligence. It's essentially a key to give edge devices with the ability to adapt to their environments and users. SO this benchmark evaluates the models capacity to successively incorporate new keywords over multiple sessions with only a handful of samples from the new classes to train with. - -This benchmark introduces a FSCIL task with streaming audio data using the MSWC (Multilingual spoken work corpus) keyword classification dataset. It is approached in 2 phases, pre-training and incremental learning. In pre-training a set of 100 words spanning 5 languages with 500 training samples each are available to train an initial model. Then for incremental learning the model goes through 10 sessions tto learn words from 10 languages in a "few-shot" learning scenario. each session adds 10 words of the corresponding language with only 5 training samples per word. after each session the model is tested in classification accuracy on all prior learned classes, Giving an evaulation on it's ability to learn new classes while retaining knowlegde about the previously learned one. Each session learns a new language resulting a knowledge base of 200 keywords by the end of the benchmark. - -* ***The following are just copy-pasted, could be messy or missing info** - * Just kept for future info - -**Event camera object detection:** Event Camera Object Detection – Object detection is a widely-used computer vision task with applications in robotics, autonomous driving, and surveillance. Such scenarios at the edge may require high energy efficiency and real-time performance, which can be achieved via event-based vision sensors[24](https://www.nature.com/articles/s41467-025-56739-4#ref-CR24 "Gallego, G. et al. Event-based vision: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 44, 154–180 (2022)."). The event camera object detection benchmark uses the Prophesee 1 Megapixel automotive detection dataset[25](https://www.nature.com/articles/s41467-025-56739-4#ref-CR25 "Perot, E., de Tournemire, P., Nitti, D., Masci, J. & Sironi, A. Learning to detect objects with a 1 megapixel event camera. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20 (2020)."), a large labeled object detection dataset with over 15 h of event camera video from the front of a car driving in various scenarios. Predetermined training, validation, and testing splits include 11.2 h, 2.2 h, and 2.2 h of recording, respectively. Pedestrian, two-wheeler, and car object classes are used in evaluation, and correctness is measured using COCO mean average precision (mAP)[26](https://www.nature.com/articles/s41467-025-56739-4#ref-CR26 "Lin, T.-Y. et al. Microsoft COCO: Common objects in context. In Computer Vision – ECCV 2014, 740–755 (2014)."). - -**Non-human Primate (NHP) Motor Prediction:** Non-human Primate (NHP) Motor Prediction – Studying models which can accurately replicate features of biological computation presents opportunities in understanding sensorimotor behavior and developing closed-loop methods for future robotic agents. It also is foundational to the development of wearable or implantable neuro-prosthetic devices that can accurately generate motor activity from neural or muscle signals. This benchmark utilizes a dataset consisting of multi-channel recordings from the sensorimotor cortex of two non-human primates (NHP Indy and NHP Loco) during reaching movements, along with corresponding fingertip motion of the reach27. Six total sessions are included from the dataset, for a total of 8712 seconds of data. The task is to train a model to predict the two-dimensional components of finger velocity using recent neural data. The sessions are treated independently (i.e., models are trained separately for each session), and the data is split to allow the first 75% for training and validation and the last 25% for evaluation. Correctness of the predictions is evaluated by the coefficient of determination (R2) score against the true finger velocity targets, averaged over all six sessions. - -**Chaotic Function Prediction:** Chaotic Function Prediction – The real-world data benchmarks presented thus far are high-dimensional and can require large networks to achieve high accuracy, raising challenges for solution types with limited I/O support and network capacity, such as mixed-signal edge prototype solutions. To address this, we include a synthetic benchmark based on prediction of one-dimensional Mackey-Glass time series28, which can be effectively tackled by smaller networks. Mackey-Glass has been widely adopted as a benchmark for evaluating temporal predictors, including neuromorphic models29,30,31. The task involves prediction of the next timestep value f(t + Δt) given the current timestep value f(t). The model is trained and validated using the first half of the time series, during which the ground truth state f(t) are supplied to the model to predict the next timestep . During the evaluation, the model uses its prior prediction to generate each next value , autoregressively forecasting the second half of the time series. Correctness is measured using symmetric mean absolute percentage error (sMAPE) of the generated time series against the target time series, a standard metric in forecasting32. The benchmark includes a set of 14 Mackey-Glass time series, which vary by the equation parameter τ, the delay constant. Lyapunov time (L), the expected predictability timescale for chaos33, is used as the time unit for each time series. The total length of each series is 20 Lyapunov times, and 75 points are sampled per Lyapunov time (Δt = L/75).
\ No newline at end of file diff --git a/Learning_notes/NeuroBench/neuroBench_notes.md b/Learning_notes/NeuroBench/neuroBench_notes.md deleted file mode 100644 index 783d32c..0000000 --- a/Learning_notes/NeuroBench/neuroBench_notes.md +++ /dev/null @@ -1,8 +0,0 @@ -## Neuromorphic computing -trying to replicate the biophysics of the brain in hardware, or "porting primitaves and computional strategies employed in the brain into engineered computing devices and algorithms". - -## difficulties in benchmarking neuromorphic computing solutions -**Lack of a formal definition**: -(left incomplete) -## Algorithm/System tracks -There are 2 different benchmark tracks aiming for faster optimization. The [[algorithm track]] is meant to evaluate in a system independant manner, meaning I don't need to worry about the implementation platform. The [[system track]] is there so I can then optimize the algorithm for the specific platform. diff --git a/Learning_notes/NeuroBench/system track.md b/Learning_notes/NeuroBench/system track.md deleted file mode 100644 index 43529de..0000000 --- a/Learning_notes/NeuroBench/system track.md +++ /dev/null @@ -1,15 +0,0 @@ -## metrics -**Correctness:** must be measured to verify the validity of the solution. No thresholds are imposed so the benchmark leaderboard must be analysed to evaluate correctness - efficiency trade offs of solutions. -**Timing:** measurements can be either sample throughput or execution time depending on the task. - -**Efficiency:** blah blah blah - -In timing and efficiency measurements both pre and post-processing must be taken into account. - -## Benchmarks -**Acoustic scene classification:** Challenges systems to sort audio into predefined categories based on the envinronmental audio context. This can allow embedded devices to adjust sound equalization profiles, target microphone denoising, and support active noise cancellation. - -Also challenges systems to fulfill technical requirements like always-on and real time operation. timing results report on device average execution time per sample. Power should be reported under idle and active contexts. Idle power measures the system prepared for inference with the model loaded, and active power measures the system running pre-processing or inference. - -**QUBO:** Quadratic unconstrained binary optimization (QUBO). basically its a series of yes or no decisions to find the best outcome. Tests the SUT using data that creates graphs based on 3 parameters, size, density, and random seeds. It runs 5 different versions of each to verify the results. After a specific amount of time neurobench cuts off the test and measures energy consumption and compares the solution to the best known solution. - @@ -1,5 +1,5 @@ # benchmark_suite_IREU -A benchmark suite for memristive and memcapacitive computing systems +A benchmark suite for memristive and memcapacitive crossbar computing systems ## overview The memristive computing field lacks standardized benchmarks that control for workload characteristics, expose diff --git a/custom_benchmark/README.md b/custom_benchmark/README.md new file mode 100644 index 0000000..7513f7f --- /dev/null +++ b/custom_benchmark/README.md @@ -0,0 +1,5 @@ +# custom benchmark for memTorch simulations + +## metrics + +## tasks diff --git a/custom_benchmark/memTorch_SNN.py b/custom_benchmark/memTorch_SNN.py new file mode 100644 index 0000000..67e9cb6 --- /dev/null +++ b/custom_benchmark/memTorch_SNN.py @@ -0,0 +1,28 @@ +import torch +import memtorch + +# 1. Standard memTorch setup (Static crossbar weights) +ann_layer = torch.nn.Linear(100, 10) +patched_layer = memtorch.mn.Module.patch_model(ann_layer, memristor_model) + +# 2. Your custom SNN wrapper loop +def forward_snn(input_spikes_over_time): + # input_spikes_over_time shape: (time_steps, batch_size, input_dim) + time_steps = input_spikes_over_time.shape[0] + v_mem = torch.zeros(batch_size, 10) # Hidden neuron membrane potentials + output_spikes = [] + + for t in range(time_steps): + # Pass binary spikes through memTorch's physical crossbar simulation + current_in = patched_layer(input_spikes_over_time[t]) + + # Leaky Integrate-and-Fire (LIF) logic (Written by you!) + v_mem = 0.9 * v_mem + current_in # Leak & Integrate + + # Fire threshold + spike = (v_mem >= 1.0).float() + v_mem[v_mem >= 1.0] = 0.0 # Reset + + output_spikes.append(spike) + + return torch.stack(output_spikes) diff --git a/custom_benchmark/memristor_benchmark.cpp b/custom_benchmark/memristor_benchmark.cpp new file mode 100644 index 0000000..91e7eb6 --- /dev/null +++ b/custom_benchmark/memristor_benchmark.cpp @@ -0,0 +1,2 @@ +#include <iostream> +#include <torch/extension.h> diff --git a/neurobench_testing/README.md b/neurobench_testing/README.md index 0e3ad72..60d3488 100644 --- a/neurobench_testing/README.md +++ b/neurobench_testing/README.md @@ -5,6 +5,7 @@ except for memtorch which needs to be cloned and compiled locally on linux syste I am not sure for windows or mac. ``` +python3 -m venv .venv pip install -r requirements.txt git clone --recursive https://github.com/coreylammie/MemTorch diff --git a/neurobench_testing/custom_memristor_model.py b/neurobench_testing/custom_memristor_model.py index 8bdcde0..eb91b99 100644 --- a/neurobench_testing/custom_memristor_model.py +++ b/neurobench_testing/custom_memristor_model.py @@ -1,5 +1,7 @@ import torch +import memtorch from memtorch.bh.memristor.Memristor import Memristor +from memtorch.utils import clip, convert_range #idk if ill need this class MemtorchMemristor(Memristor): def __init__( @@ -27,16 +29,8 @@ class MemtorchMemristor(Memristor): self.i_off = i_off self.p = p - # makes sure w starts in valid state - if not hasattr(self, 'w'): - self.w = torch.tensor(0.5) - - - """ - Updates w and computes new resistance - - """ - def step(self, v, dt): - i = v / self.r_curr - + #state variables + self.w = 0.5 + self.g = 1/self.r_on + def window() diff --git a/neurobench_testing/kaevin_memristor.py b/neurobench_testing/kaevin_memristor.py new file mode 100644 index 0000000..b4a4364 --- /dev/null +++ b/neurobench_testing/kaevin_memristor.py @@ -0,0 +1,134 @@ +import numpy as np +from scipy.integrate import solve_ivp + +# TEAM (Threshold Adaptive Memristor) model as a reusable class +class TEAMMemristor: + + def __init__( + self, + k_off=1, # switching rate for off-state + k_on=-1, # switching rate for on-state + alpha_off=5, # exponent controlling nonlinearity when switching off + alpha_on=5, # exponent controlling nonlinearity when switching on + i_off=0.5e-3, # threshold current to trigger off state switching + i_on=-0.5e-3, # threshold current to trigger on state switching + g_on=1/1e3, # maximum conductance (1 kohm = 1 ms) + g_off=1/10e3, # minimum conductance (10 kohm = 0.1 ms) + w_init=0.5, # initial state variable (0=off, 1=on) + p=2 # window function exponent + ): + self.k_off = k_off + self.k_on = k_on + self.alpha_off = alpha_off + self.alpha_on = alpha_on + self.i_off = i_off + self.i_on = i_on + self.G_on = G_on + self.G_off = G_off + self.w_init = w_init + self.p = p + + def set_state(self, w): + # Update initial condition for next simulation + self.w_init = np.clip(w, 0, 1) + + def window(self, w, i): + # Nonlinear window function: reduces switching rate near the boundaries + w = np.clip(w, 0.0, 1.0) + if i >= 0: + return 1 - w**(2*self.p) # switching off: slower near w=1 + else: + return 1 - (1-w)**(2*self.p) # switching on: slower near w=0 + + def conductance(self, w): + # Linear interpolation between off and on conductance based on state w + w = np.clip(w, 0.0, 1.0) + return self.G_off + w*(self.G_on - self.G_off) + + def dw_dt(self, w, i): + # TEAM state dynamics: dw/dt depends on current magnitude and direction + w = np.clip(w, 0.0, 1.0) + + if i >= self.i_off: # positive current above threshold = switch off + dw = ( + self.k_off + * ((i/self.i_off)-1)**self.alpha_off + * self.window(w, i) + ) + elif i <= self.i_on: # negative current below threshold = switch on + dw = ( + self.k_on + * (((-i)/abs(self.i_on))-1)**self.alpha_on + * self.window(w, i) + ) + else: # between thresholds = no switching + dw = 0.0 + + # Enforce physical bounds: prevent state from leaving [0,1] + if w <= 0 and dw < 0: + dw = 0 + if w >= 1 and dw > 0: + dw = 0 + + return dw + + def simulate(self, + freq=1, # excitation frequency (Hz) + V_amp=1.5, # sinusoid amplitude (V) + cycles=3): # number of periods to simulate + # Solve the memristor ODE for given frequency and voltage amplitude + + def voltage(t): # sinusoidal excitation signal + return V_amp*np.sin(2*np.pi*freq*t) + + def ode(t, y): # dy/dt: current through memristor + w = y[0] + v = voltage(t) + G = self.conductance(w) + i = G*v # Ohm's law: i = G*v + return [self.dw_dt(w, i)] + + T = 1/freq # period + t_end = cycles*T + t_eval = np.linspace(0, t_end, 10000) # dense time grid for smooth curves + + # Solve with RK45, small max step + sol = solve_ivp( + ode, + [0, t_end], + [self.w_init], + t_eval=t_eval, + method='RK45', + max_step=T/1000, # max step keeps resolution within one period + rtol=1e-8, + atol=1e-10 + ) + + raw_w = sol.y[0] + + # Check if numerical solver violated physical bounds + eps = 1e-6 + if np.any(raw_w < -eps) or np.any(raw_w > 1+eps): + print( + "WARNING: solver left bounds " + f"min={raw_w.min():.12f}, " + f"max={raw_w.max():.12f}" + ) + + w = np.clip(raw_w, 0, 1) # enforce bounds just in case + t = sol.t + v = voltage(t) + G = self.conductance(w) + i = G*v # current throughout simulation + + return t, w, v, i + + def resistance(self, w): + # Compute resistance as reciprocal of conductance + return 1/self.conductance(w) + + def reset(self): + # Reset state to default initial condition + self.w_init = 0.5 + + diff --git a/neurobench_testing/memTorch_testing.py b/neurobench_testing/memTorch_testing.py index 5c3d828..1992e05 100644 --- a/neurobench_testing/memTorch_testing.py +++ b/neurobench_testing/memTorch_testing.py @@ -1,15 +1,10 @@ -""" -This is a program to test out using memTorch with neurobench. -June 17th, 2026 -Author: Tanner Robison, -Teuscher Lab -""" +import sys import torch import torch.nn as nn import snntorch as snn from snntorch import surrogate -from torch.utils.data import DataLoader +from torch.utils.data import DataLoader, Subset from neurobench.models import SNNTorchModel from neurobench.benchmarks import Benchmark, benchmark @@ -32,31 +27,9 @@ import copy from memtorch.mn.Module import patch_model from memtorch.map.Parameter import naive_map from memtorch.bh.memristor import VTEAM +from memtorch.map.Input import naive_scale beta = 0.9 -class SNN(nn.Module): - def __init__(self): - super().__init__() - - #standard layers - self.fc1 = nn.Linear(20, 128) - self.fc2 = nn.Linear(128, 35) - - #Spiking neurons - self.lif1 = snn.Leaky(beta=beta, init_hidden=True) - self.lif2 = snn.Leaky(beta=beta, init_hidden=True, output=True) - - def forward(self, x): - x = x.view(x.size(0), -1) - - cur1 = self.fc1(x) - spk1 = self.lif1(cur1) - - cur2 = self.fc2(spk1) - spk2, mem2 = self.lif2(cur2) - - return spk2, mem2 - device = torch.device("cpu") spike_grad = surrogate.fast_sigmoid() net = nn.Sequential( @@ -71,28 +44,38 @@ net = nn.Sequential( snn.Leaky(beta=beta, spike_grad=spike_grad, init_hidden=True, output=True), ) -#memristor patch -reference_memristor = VTEAM() +vteam_params = { + 'time_series_resolution': 1e-3, + 'r_on': 50, + 'r_off': 1000, +} net.load_state_dict(torch.load("examples/gsc/model_data/s2s_gsc_snntorch", map_location=device)) patched_net = patch_model( copy.deepcopy(net), - memristor_model=reference_memristor, - memristor_model_params={'time_series_resolution': 1e-8}, + memristor_model=VTEAM, + memristor_model_params=vteam_params, + module_parameters_to_patch=[torch.nn.Linear], mapping_routine=naive_map, transistor=True, tile_shape=(128, 128), - ADC_resolution=8, - use_bindings=True + max_input_voltage=0.3, + scaling_routine=naive_scale, + ADC_resolution=16, + use_bindings=True, + verbose=True, ) static_metrics = [Footprint, ConnectionSparsity] workload_metrics = [ActivationSparsity, SynapticOperations, ClassificationAccuracy] -# data loader here maybe?? test_set = SpeechCommands(path="data/SpeechCommands/", subset="testing") -test_set_loader = DataLoader(test_set, batch_size=500, shuffle=True) + +#shorten data set so I can actually run it lol +tiny_indices = list(range(10)) +tiny_test_set = Subset(test_set, tiny_indices) +test_set_loader = DataLoader(tiny_test_set, batch_size=16, shuffle=True) pre_processor = [S2SPreProcessor(device=device)] post_processor = [ChooseMaxCount()] @@ -106,13 +89,36 @@ benchmark = Benchmark( [static_metrics, workload_metrics] ) +print("\n --Checking signal strength--") +dummy_input = torch.randn(2, 20).to(device) + +try: + raw_signal = patched_net(dummy_input) + + if isinstance(raw_signal, tuple) and len(raw_signal) > 1: + voltage_signal = raw_signal[0] + print("Checking Neuron voltages") + else: + voltage_signal = raw_signal + print("Checking RAW OUTPUTS") + + + print(f"Signal Max: {raw_signal.max().item():.8f}") + print(f"Signal Min: {raw_signal.min().item():.8f}") + print(f"Signal Mean: {raw_signal.mean().item():.8f}") + +except Exception as e: + print("Error getting signal:", e) + +sys.exit() + results = benchmark.run() print("\n\n----- IDEAL BENCHMARK -----") print(f"Footprint: {results['Footprint']}") print(f"Connection Sparsity: {results['ConnectionSparsity']}") print(f"Activation Sparsity: {results['ActivationSparsity']}") print(f"Synaptic Operations: {results['SynapticOperations']}") -print(f"Classification Accuracy: {results['ClassificationAccuracy']}") +print(f"Classification Accuracy: {results['ClassificationAccuracy']}\n") #energy calculations ENERGY_PER_MAC = 0.9e-12 @@ -125,7 +131,7 @@ total_energy = (macs * ENERGY_PER_MAC) + (acs * ENERGY_PER_AC) print("Energy Report:") print(f"Total Operations: {macs} MACS, {acs} acs ") -print(f"Calculated Energy cost: {total_energy} joules per batch") +print(f"Calculated Energy cost: {total_energy} joules per batch\n\n") model = SNNTorchModel(patched_net) @@ -137,8 +143,9 @@ benchmark = Benchmark( [static_metrics, workload_metrics] ) + with torch.no_grad(): #makes sure its in inference mode - #otherwise you get memory leaks + #otherwise you get memory leaks : ( results = benchmark.run() print("\n\n----- MEMRISTOR BENCHMARK -----") diff --git a/neurobench_testing/memTorch_testing_custom_memristor.py b/neurobench_testing/memTorch_testing_custom_memristor.py deleted file mode 100644 index 64f338e..0000000 --- a/neurobench_testing/memTorch_testing_custom_memristor.py +++ /dev/null @@ -1,160 +0,0 @@ -from memtorch import memristor -import torch -import torch.nn as nn -import snntorch as snn -from snntorch import surrogate - -from torch.utils.data import DataLoader - -from neurobench.models import SNNTorchModel -from neurobench.benchmarks import Benchmark, benchmark -from neurobench.datasets import SpeechCommands - -from neurobench.metrics.workload import ( - ActivationSparsity, - SynapticOperations, - ClassificationAccuracy, -) - -from neurobench.metrics.static import ( - Footprint, - ConnectionSparsity, -) - -from neurobench.processors.preprocessors import S2SPreProcessor -from neurobench.processors.postprocessors import ChooseMaxCount - -import copy -from memtorch.mn.Module import patch_model -from memtorch.map.Parameter import naive_map -from memtorch.bh.memristor import VTEAM - -from kaevin_memristor import TEAMMemristor - -beta = 0.9 -class SNN(nn.Module): - def __init__(self): - super().__init__() - - #standard layers - self.fc1 = nn.Linear(20, 128) - self.fc2 = nn.Linear(128, 35) - - #Spiking neurons - self.lif1 = snn.Leaky(beta=beta, init_hidden=True) - self.lif2 = snn.Leaky(beta=beta, init_hidden=True, output=True) - - def forward(self, x): - x = x.view(x.size(0), -1) - - cur1 = self.fc1(x) - spk1 = self.lif1(cur1) - - cur2 = self.fc2(spk1) - spk2, mem2 = self.lif2(cur2) - - return spk2, mem2 - -device = torch.device("cpu") -spike_grad = surrogate.fast_sigmoid() -net = nn.Sequential( - nn.Flatten(), - nn.Linear(20, 256), - snn.Leaky(beta=beta, spike_grad=spike_grad, init_hidden=True), - nn.Linear(256, 256), - snn.Leaky(beta=beta, spike_grad=spike_grad, init_hidden=True), - nn.Linear(256, 256), - snn.Leaky(beta=beta, spike_grad=spike_grad, init_hidden=True), - nn.Linear(256, 35), - snn.Leaky(beta=beta, spike_grad=spike_grad, init_hidden=True, output=True), -) - -#memristor patch -reference_memristor = TEAMMemristor - -net.load_state_dict(torch.load("examples/gsc/model_data/s2s_gsc_snntorch", map_location=device)) - -patched_net = patch_model( - copy.deepcopy(net), - memristor_model=reference_memristor, - memristor_model_params={'time_series_resolution': 1e-8}, - mapping_routine=naive_map, - transistor=True, - tile_shape=(128, 128), - ADC_resolution=8, - use_bindings=True -) - -static_metrics = [Footprint, ConnectionSparsity] -workload_metrics = [ActivationSparsity, SynapticOperations, ClassificationAccuracy] - -# data loader here maybe?? -test_set = SpeechCommands(path="data/SpeechCommands/", subset="testing") -test_set_loader = DataLoader(test_set, batch_size=500, shuffle=True) - -pre_processor = [S2SPreProcessor(device=device)] -post_processor = [ChooseMaxCount()] - -model = SNNTorchModel(net) -benchmark = Benchmark( - model, - test_set_loader, - pre_processor, - post_processor, - [static_metrics, workload_metrics] -) - -results = benchmark.run() -print("\n\n----- IDEAL BENCHMARK -----") -print(f"Footprint: {results['Footprint']}") -print(f"Connection Sparsity: {results['ConnectionSparsity']}") -print(f"Activation Sparsity: {results['ActivationSparsity']}") -print(f"Synaptic Operations: {results['SynapticOperations']}") -print(f"Classification Accuracy: {results['ClassificationAccuracy']}") - -#energy calculations -ENERGY_PER_MAC = 0.9e-12 -ENERGY_PER_AC = 0.1e-12 - -macs = results['SynapticOperations']['Effective_MACs'] -acs = results['SynapticOperations']['Effective_ACs'] - -total_energy = (macs * ENERGY_PER_MAC) + (acs * ENERGY_PER_AC) - -print("Energy Report:") -print(f"Total Operations: {macs} MACS, {acs} acs ") -print(f"Calculated Energy cost: {total_energy} joules per batch") - - -model = SNNTorchModel(patched_net) -benchmark = Benchmark( - model, - test_set_loader, - pre_processor, - post_processor, - [static_metrics, workload_metrics] -) - -with torch.no_grad(): #makes sure its in inference mode - #otherwise you get memory leaks - results = benchmark.run() - -print("\n\n----- MEMRISTOR BENCHMARK -----") -print(f"Footprint: {results['Footprint']}") -print(f"Connection Sparsity: {results['ConnectionSparsity']}") -print(f"Activation Sparsity: {results['ActivationSparsity']}") -print(f"Synaptic Operations: {results['SynapticOperations']}") -print(f"Classification Accuracy: {results['ClassificationAccuracy']}") - -#energy calculations -ENERGY_PER_MAC = 0.9e-12 -ENERGY_PER_AC = 0.1e-12 - -macs = results['SynapticOperations']['Effective_MACs'] -acs = results['SynapticOperations']['Effective_ACs'] - -total_energy = (macs * ENERGY_PER_MAC) + (acs * ENERGY_PER_AC) - -print("Energy Report:") -print(f"Total Operations: {macs} MACS, {acs} acs ") -print(f"Calculated Energy cost: {total_energy} joules per batch") diff --git a/quickstart b/quickstart Binary files differnew file mode 100755 index 0000000..1c49674 --- /dev/null +++ b/quickstart diff --git a/references/kaevin_memristor.py b/references/kaevin_memristor.py index b4a4364..f186958 100644 --- a/references/kaevin_memristor.py +++ b/references/kaevin_memristor.py @@ -23,8 +23,8 @@ class TEAMMemristor: self.alpha_on = alpha_on self.i_off = i_off self.i_on = i_on - self.G_on = G_on - self.G_off = G_off + self.g_on = g_on + self.g_off = g_off self.w_init = w_init self.p = p diff --git a/spires b/spires new file mode 160000 +Subproject c9544bdcd3fe9be7d055a3e1e120110ef4fe258 diff --git a/spires_tutorial.c b/spires_tutorial.c new file mode 100644 index 0000000..758b5ea --- /dev/null +++ b/spires_tutorial.c @@ -0,0 +1,79 @@ +#include <../src/neurons/lif_discrete.h> +#include <math.h> +#include <spires.h> +#include <stdio.h> +#include <stdlib.h> + +#define N_TRAIN 500 +#define N_TEST 100 +#define PI 3.14159265358979323846 + +int main(void) { + /* 1. Configure the reservoir */ + double lif_cfg[] = {0.0, 1.0, 0.2, 0.5}; + + spires_reservoir_config cfg = { + .num_neurons = 400, + .num_inputs = 1, + .num_outputs = 1, + .spectral_radius = 0.95, + .ei_ratio = 0.8, + .input_strength = 0.1, + .connectivity = 0.1, + .dt = 1.0, + .connectivity_type = SPIRES_CONN_RANDOM, + .neuron_type = SPIRES_NEURON_LIF_DISCRETE, + .neuron_params = lif_cfg, + }; + + /* 2. Create the reservoir */ + spires_reservoir *r = NULL; + spires_status s = spires_reservoir_create(&cfg, &r); + if (s != SPIRES_OK) { + fprintf(stderr, "Failed to create reservoir: %d\n", s); + return 1; + } + + /* 3. Generate training data — predict sin(t+1) from sin(t) */ + double input_train[N_TRAIN]; + double target_train[N_TRAIN]; + for (int i = 0; i < N_TRAIN; i++) { + input_train[i] = sin(2.0 * PI * i / 50.0); + target_train[i] = sin(2.0 * PI * (i + 1) / 50.0); + } + + /* 4. Train with ridge regression */ + s = spires_train_ridge(r, input_train, target_train, N_TRAIN, 1e-6); + if (s != SPIRES_OK) { + fprintf(stderr, "Training failed: %d\n", s); + spires_reservoir_destroy(r); + return 1; + } + + /* 5. Generate test input */ + double input_test[N_TEST]; + for (int i = 0; i < N_TEST; i++) { + input_test[i] = sin(2.0 * PI * (N_TRAIN + i) / 50.0); + } + + /* 6. Run inference */ + double *predictions = spires_run(r, input_test, N_TEST); + if (!predictions) { + fprintf(stderr, "Inference failed\n"); + spires_reservoir_destroy(r); + return 1; + } + + /* 7. Print a few predictions vs. expected values */ + printf("Step | Predicted | Expected\n"); + printf("-----+-----------+---------\n"); + for (int i = 0; i < 10; i++) { + double expected = sin(2.0 * PI * (N_TRAIN + i + 1) / 50.0); + printf("%4d | %+.5f | %+.5f\n", i, predictions[i], expected); + } + + /* 8. Clean up */ + free(predictions); + spires_reservoir_destroy(r); + return 0; +} |
