summaryrefslogtreecommitdiff
path: root/Learning_notes/NeuroBench
diff options
context:
space:
mode:
Diffstat (limited to 'Learning_notes/NeuroBench')
-rw-r--r--Learning_notes/NeuroBench/algorithm track.md25
-rw-r--r--Learning_notes/NeuroBench/neuroBench_notes.md8
-rw-r--r--Learning_notes/NeuroBench/system track.md15
3 files changed, 0 insertions, 48 deletions
diff --git a/Learning_notes/NeuroBench/algorithm track.md b/Learning_notes/NeuroBench/algorithm track.md
deleted file mode 100644
index aaaedc8..0000000
--- a/Learning_notes/NeuroBench/algorithm track.md
+++ /dev/null
@@ -1,25 +0,0 @@
-aims to evaluate algorithms in a system independant manner. This means that the implementation platform can be ill matched to the algorithm benchmark it executes. Which allows for the algorithm complexity and expected performance to be examined in a theoretical manner, resulting in quicker prototyping and and functional analysis.
-
-## metrics
-**footprint:** A measure of the memory footprint in bytes. This metric summarizes synaptic weight count, weight precision, trainable neuron paramaters, data buffers, etc... Zero weight are included because they are distinguished in the connection sparsity metric.
-
-**Connection sparsity:** is the the number of zero weights divided by the total number of weights, accumulated over all layers. 0 refers to no sparsity (fully connected) and 1 means no sparsity.
-
-**Activation sparsity:** The average sparsity of neuron activations over all neurons in all model layers, for all timesteps of all tested samples. again 0 refers to no sparsity (all neurons are activated) and 1 refers to the case where all neurons have a zero output.
-
-**Synaptic Operations:** Average number of synaptic operation per model execution, based on neuron activations and the associated fanout synapses. The metric can be subdivided into dense, effective multiply-accumulate, and effective accumulate synaptic operations (dense, Eff_MACs, Eff_ACs). Dense accounts for all zero and non-zero neuron activations and synaptic connections, and reflects the number of operations necessary on hardware that does not support sparsity. Eff_MACs and Eff_ACs only count effective synaptic operations by disregarding zero activations and zero connections, thus reflecting operation cost on sparsity-aware hardware. Synaptic operations with non-binary activation are considered multiply accumulates(Eff_MACs), while those with binary activation are considered accumulates (ACs).
-
-Footprint & connection sparsity are static metrics, so they can be determined from the model alone. activation sparsity, synaptic operations, and correctness are classified as workload metrics, so they depend on the execution of the model based on the benchmark data.
-## benchmarks
-**FSCIL:** Few-Shot Class Incremental Learning (FSCIL) is learning new tasks from a small amount of experiences while retaining knowledge of prior tasks is a key characteristic of biological intelligence. It's essentially a key to give edge devices with the ability to adapt to their environments and users. SO this benchmark evaluates the models capacity to successively incorporate new keywords over multiple sessions with only a handful of samples from the new classes to train with.
-
-This benchmark introduces a FSCIL task with streaming audio data using the MSWC (Multilingual spoken work corpus) keyword classification dataset. It is approached in 2 phases, pre-training and incremental learning. In pre-training a set of 100 words spanning 5 languages with 500 training samples each are available to train an initial model. Then for incremental learning the model goes through 10 sessions tto learn words from 10 languages in a "few-shot" learning scenario. each session adds 10 words of the corresponding language with only 5 training samples per word. after each session the model is tested in classification accuracy on all prior learned classes, Giving an evaulation on it's ability to learn new classes while retaining knowlegde about the previously learned one. Each session learns a new language resulting a knowledge base of 200 keywords by the end of the benchmark.
-
-* ***The following are just copy-pasted, could be messy or missing info**
- * Just kept for future info
-
-**Event camera object detection:** Event Camera Object Detection – Object detection is a widely-used computer vision task with applications in robotics, autonomous driving, and surveillance. Such scenarios at the edge may require high energy efficiency and real-time performance, which can be achieved via event-based vision sensors[24](https://www.nature.com/articles/s41467-025-56739-4#ref-CR24 "Gallego, G. et al. Event-based vision: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 44, 154–180 (2022)."). The event camera object detection benchmark uses the Prophesee 1 Megapixel automotive detection dataset[25](https://www.nature.com/articles/s41467-025-56739-4#ref-CR25 "Perot, E., de Tournemire, P., Nitti, D., Masci, J. & Sironi, A. Learning to detect objects with a 1 megapixel event camera. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20 (2020)."), a large labeled object detection dataset with over 15 h of event camera video from the front of a car driving in various scenarios. Predetermined training, validation, and testing splits include 11.2 h, 2.2 h, and 2.2 h of recording, respectively. Pedestrian, two-wheeler, and car object classes are used in evaluation, and correctness is measured using COCO mean average precision (mAP)[26](https://www.nature.com/articles/s41467-025-56739-4#ref-CR26 "Lin, T.-Y. et al. Microsoft COCO: Common objects in context. In Computer Vision – ECCV 2014, 740–755 (2014).").
-
-**Non-human Primate (NHP) Motor Prediction:** Non-human Primate (NHP) Motor Prediction – Studying models which can accurately replicate features of biological computation presents opportunities in understanding sensorimotor behavior and developing closed-loop methods for future robotic agents. It also is foundational to the development of wearable or implantable neuro-prosthetic devices that can accurately generate motor activity from neural or muscle signals. This benchmark utilizes a dataset consisting of multi-channel recordings from the sensorimotor cortex of two non-human primates (NHP Indy and NHP Loco) during reaching movements, along with corresponding fingertip motion of the reach27. Six total sessions are included from the dataset, for a total of 8712 seconds of data. The task is to train a model to predict the two-dimensional components of finger velocity using recent neural data. The sessions are treated independently (i.e., models are trained separately for each session), and the data is split to allow the first 75% for training and validation and the last 25% for evaluation. Correctness of the predictions is evaluated by the coefficient of determination (R2) score against the true finger velocity targets, averaged over all six sessions.
-
-**Chaotic Function Prediction:** Chaotic Function Prediction – The real-world data benchmarks presented thus far are high-dimensional and can require large networks to achieve high accuracy, raising challenges for solution types with limited I/O support and network capacity, such as mixed-signal edge prototype solutions. To address this, we include a synthetic benchmark based on prediction of one-dimensional Mackey-Glass time series28, which can be effectively tackled by smaller networks. Mackey-Glass has been widely adopted as a benchmark for evaluating temporal predictors, including neuromorphic models29,30,31. The task involves prediction of the next timestep value f(t + Δt) given the current timestep value f(t). The model is trained and validated using the first half of the time series, during which the ground truth state f(t) are supplied to the model to predict the next timestep . During the evaluation, the model uses its prior prediction to generate each next value , autoregressively forecasting the second half of the time series. Correctness is measured using symmetric mean absolute percentage error (sMAPE) of the generated time series against the target time series, a standard metric in forecasting32. The benchmark includes a set of 14 Mackey-Glass time series, which vary by the equation parameter τ, the delay constant. Lyapunov time (L), the expected predictability timescale for chaos33, is used as the time unit for each time series. The total length of each series is 20 Lyapunov times, and 75 points are sampled per Lyapunov time (Δt = L/75). \ No newline at end of file
diff --git a/Learning_notes/NeuroBench/neuroBench_notes.md b/Learning_notes/NeuroBench/neuroBench_notes.md
deleted file mode 100644
index 783d32c..0000000
--- a/Learning_notes/NeuroBench/neuroBench_notes.md
+++ /dev/null
@@ -1,8 +0,0 @@
-## Neuromorphic computing
-trying to replicate the biophysics of the brain in hardware, or "porting primitaves and computional strategies employed in the brain into engineered computing devices and algorithms".
-
-## difficulties in benchmarking neuromorphic computing solutions
-**Lack of a formal definition**:
-(left incomplete)
-## Algorithm/System tracks
-There are 2 different benchmark tracks aiming for faster optimization. The [[algorithm track]] is meant to evaluate in a system independant manner, meaning I don't need to worry about the implementation platform. The [[system track]] is there so I can then optimize the algorithm for the specific platform.
diff --git a/Learning_notes/NeuroBench/system track.md b/Learning_notes/NeuroBench/system track.md
deleted file mode 100644
index 43529de..0000000
--- a/Learning_notes/NeuroBench/system track.md
+++ /dev/null
@@ -1,15 +0,0 @@
-## metrics
-**Correctness:** must be measured to verify the validity of the solution. No thresholds are imposed so the benchmark leaderboard must be analysed to evaluate correctness - efficiency trade offs of solutions.
-**Timing:** measurements can be either sample throughput or execution time depending on the task.
-
-**Efficiency:** blah blah blah
-
-In timing and efficiency measurements both pre and post-processing must be taken into account.
-
-## Benchmarks
-**Acoustic scene classification:** Challenges systems to sort audio into predefined categories based on the envinronmental audio context. This can allow embedded devices to adjust sound equalization profiles, target microphone denoising, and support active noise cancellation.
-
-Also challenges systems to fulfill technical requirements like always-on and real time operation. timing results report on device average execution time per sample. Power should be reported under idle and active contexts. Idle power measures the system prepared for inference with the model loaded, and active power measures the system running pre-processing or inference.
-
-**QUBO:** Quadratic unconstrained binary optimization (QUBO). basically its a series of yes or no decisions to find the best outcome. Tests the SUT using data that creates graphs based on 3 parameters, size, density, and random seeds. It runs 5 different versions of each to verify the results. After a specific amount of time neurobench cuts off the test and measures energy consumption and compares the solution to the best known solution.
-