Darknet Framework for Machine Learning Developed by Joseph
Darknet Framework for Machine Learning Developed by Joseph Redmon https: //pjreddie. com/darknet/ Presented by Alan Mishchenko UC Berkeley
Overview l Introducing machine learning (ML) l l Implementation of ML l l Getting started, learning from the code Case studies l l l Data flow during training and evaluation The computation challenge Software and hardware implementation Darknet: A complete open-source ML solution l l Artificial neuron, neural network, convolution, etc Tiny Darknet: Image classification YOLO: Real-time object detection Further improving Darknet framework l l Speeding up computation Adding quantization
Machine Learning (ML) l ML learns information from data, for example l l l Data is composed of data instances l l One image is a data instance Data instances are divided into two groups: l l l Classify the image (dog, cat, etc) Detect where an object is located in the image Training set is used for training Validation set is used to evaluate accuracy ML model is one specific way to do ML l Neural networks, random forests, etc
Accuracy of ML Models l l In a typical ML scenario, training data is prepared and used to train an ML model in several iterations l The more training, the better the result (hopefully) A trained ML model takes an input data sample and produces the result of classification (correct or incorrect) Accuracy is determined by counting the percentage of correct answers A typical learning curve looks as follows:
Neural Networks l l Neural networks are widely used in machine learning Network is composed of layers l l Each layer includes neurons and/or other operators A typical use of neural networks is in image recognition l l l The input data is an image The output data is the result of classification, detection, etc The architecture determines the quality/runtime tradeoff
Neural Network https: //towardsdatascience. com/converting-a-simple-deep-learning-model-from-pytorch-to-tensorflow-b 6 b 353351 f 5 d
Neuron https: //medium. com/@jayeshbahire/the-artificial-neural-networks-handbook-part-4 -d 2087 d 1 f 583 e
Convolution https: //datascience. stackexchange. com/questions/23183/why-convolutions-always-use-odd-numbers-as-filter-size
Max Pooling and Average Pooling https: //www. researchgate. net/figure/Illustration-of-Max-Pooling-and-Average-Pooling-Figure-2 -above-shows-an-example-of-max_fig 2_333593451
A Complete Neural Network (VGG-16) https: //neurohive. io/en/popular-networks/vgg 16/
Res. Net K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition”, Microsoft Research, Dec 2015, https: //arxiv. org/pdf/1512. 03385. pdf
Several Versions of Res. Net A. Sachan, “Detailed guide to understand implement Res. Nets”. https: //cv-tricks. com/keras/understand-implement-resnets/
Overview l Introducing machine learning (ML) l l Implementation of ML l l Getting started, learning from the code Case studies l l l Data flow during training and evaluation The computation challenge of neural networks Software and hardware implementation Darknet: A complete open-source ML solution l l Artificial neuron, neural network, convolution, etc Tiny Darknet: Image classification YOLO: Real-time object detection Further improving Darknet framework l l Speeding up computation Adding quantization
Machine Learning Flow l Inference: Evaluate network on one data instance (one run) Training: Compute network parameters on many instances (many runs) l CNN is a sequence of convolutional layers having l l Input and output data (i, o), Convolution parameters (w, b), Gradients ( ) Forward pass (inference, training) w i Convolution o = i * w + b (convolution) o b Backward pass (training) w w = i * o (convolution) i = o * w’ (convolution) b = o (addition) i Convolution b https: //towardsdatascience. com/backpropagation-in-a-convolutional-layer-24 c 8 d 64 d 8509 o
Neural Network Training l l Prepare training data (each image has input data and output class) For a group of images called mini-batch (~100 images) l l l Calculate the output for these images using the forward pass Calculate the error (the difference of the actual and the expected output) Calculate the gradients using the backward pass Update the weights by subtracting a fraction of the gradient Iterate the above procedure ~100 times for all images l or until the accuracy is good (the error is small) Calculate error Forward pass Start with random weights Stop when error is small Backward pass Update weights
ML Challenges l Making it work l l Is learning taking place during training? Does accuracy continue to improve? Is there overfitting? Being able to compute l l Inference Training
The Computation Challenge l Consider a medium-sized neural network, Res. Net 18 l Inference for one image (one forward pass) l l Forward and backward pass for one image l l l Backward pass = ~2 x compute of forward pass 3 * 3. 6 GFLOP = ~10 GFLOP Training on Image. Net (~1 M images * ~100 epochs) l l 3. 6 GFLOP (3. 6*10^9 floating-point operations) 100 M images * (forward pass + backward pass) 100 M * 10 GFLOP = 10^18 FLOP (1 EFLOP) ~300 M times slower than inference for one image Summary: l l Inference = ~3. 6 GFLOP Training = ~10^9 GFLOP
Runtime on a CPU l What is the inference runtime of Res. Net 18 (3. 6 GFLOP) on a CPU? l l l (I tried it on my computer while preparing these slides) Running on one core, Darknet takes 1 sec to evaluate Res. Net 18 for one image How long it would have taken to train Res. Net 18 on Image. Net? l 100 M times longer than inference for one image: 3. 1 years l The computer’s performance is 3. 6 GFLOP/sec l How far is it from the peak performance of this computer? l The peak performance of this computer is (see the next slide) l l l The performance of a typical GPU is ~10, 000 GFLOP/sec l l 40 GFLOP/sec (1 core) (my computer is 9 years old; a new is ~4 x faster: 160 GFLOP/sec) 1, 280 GFLOP/sec (32 cores) Interesting: GPU is faster than multi-core CPU, but no more than 10 x! What would be the runtime, if my computer had peak performance? l l l 10 x faster on 1 core (inference: 0. 1 sec) 320 x faster on 32 cores (training: 9 days) Training on GPU would have taken 1. 1 days
Peak Performance of a CPU 32 -bit FLOP/Cycle for Intel CPUs (https: //en. wikipedia. org/wiki/FLOPS) Microarchitecture ISA FP 32 Intel Atom (Bonnell, Saltwell, Silvermont and Goldmont) SSE 3 (64 -bit) 4 Intel Core (Merom, Penryn), Nehalem (Nehalem, Westmere) SSE 4 (128 -bit) 8 Intel Sandy Bridge (Sandy Bridge, Ivy Bridge) AVX (256 -bit) 16 Intel Haswell (Haswell, Devil's Canyon, Broadwell), Skylake AVX 2 & FMA (256 -bit) 32 Intel Xeon Phi (Knights Corner) SSE & FMA (256 -bit) 32 Intel Skylake-X, Xeon Phi (Knights Landing, Knights Mill) AVX-512 & FMA (512 -bit) 64 [user@computer] ~> lscpu Architecture: Model name: x 86_64 Intel(R) Xeon(R) CPU E 5 -2670 0 @ 2. 60 GHz CPU(s): 32 Flags: fpu vme de pse tsc msr pae mce cx 8 apic sep mtrr pge mca cmov pat pse 36 clflush dts acpi mmx fxsr sse 2 ss ht tm pbe syscall nx pdpe 1 gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf pni pclmulqdq dtes 64 monitor ds_cpl vmx smx est tm 2 ssse 3 cx 16 xtpr pdcm pcid dca sse 4_1 sse 4_2 x 2 apic popcnt tsc_deadline_timer aes xsave avx lahf_lm epb ssbd ibrs ibpb stibp kaiser tpr_shadow vnmi flexpriority ept vpid xsaveopt dtherm ida arat pln pts flush_l 1 d Peak Performance = 2. 6 GHz * 16 FLOP/cycle = ~40 GFLOP/sec (1 core) = ~1280 GFLOP/sec (32 cores)
Software and Hardware Implementation l Training l l l Currently performed in software Can be made much faster using GPUs Inference l l Can be performed in software There is a growing demand to perform it in hardware l One of the reasons why quantization is important
Overview l Introducing machine learning (ML) l l Implementation of ML l l Getting started, learning from the code Case studies l l l Data flow during training and evaluation The computation challenge Software and hardware implementation Darknet: A complete open-source ML solution l l Artificial neuron, neural network, convolution, etc Tiny Darknet: Image classification YOLO: Real-time object detection Further improving Darknet framework l l Speeding up computation Adding quantization
Darknet: An Open-Source ML Solution
Features of Darknet l A complete ML framework with clean, self-contained code for l l There are several pre-trained neural networks l l Image processing, training, validation, inference, etc Can be used immediately for image classification, object detection, etc The code can be studied and customized for related projects l For example, this is the forward pass through the convolution layer: https: //github. com/pjreddie/darknet/blob/master/src/convolutional_layer. c
Overview l Introducing machine learning (ML) l l Implementation of ML l l Getting started, learning from the code Case studies l l l Data flow during training and evaluation The computation challenge Software and hardware implementation Darknet: A complete open-source ML solution l l Artificial neuron, neural network, convolution, etc Tiny Darknet: Image classification YOLO: Real-time object detection Further improving Darknet framework l l Speeding up computation Adding quantization
Tiny Darknet l A lightweight image classifier https: //pjreddie. com/darknet/tiny-darknet/
Classification Using Pre-Trained Tiny Darknet Install Darknet: git clone https: //github. com/pjreddie/darknet cd darknet make Download pretrained weights: wget https: //pjreddie. com/media/files/tiny. weights Run the classifier: darknet classify cfg/tiny. cfg tiny. weights data/dog. jpg The output will look as follows: data/dog. jpg: Predicted in 0. 16 seconds. malamute: 0. 167168 Eskimo dog: 0. 065828 dogsled: 0. 063020 standard schnauzer: 0. 051153 Siberian husky: 0. 037506
YOLO: Real-Time Object Detection l YOLO is designed to detect multiple objects in one image at different scales
YOLO: Illustration
Detection Using Pre-Trained YOLO Install Darknet: git clone https: //github. com/pjreddie/darknet cd darknet make Download pretrained weights: wget https: //pjreddie. com/media/files/yolov 3. weights Run the detector: darknet detect cfg/yolov 3. cfg yolov 3. weights data/dog. jpg The output will look as follows: layer filters size input output 0 conv 32 3 x 3 / 1 416 x 3 -> 416 x 32 0. 299 BFLOPs 1 conv 64 3 x 3 / 2 416 x 32 -> 208 x 64 1. 595 BFLOPs. . . . 105 conv 255 1 x 1 / 1 52 x 256 -> 52 x 255 0. 353 BFLOPs 106 detection truth_thresh: Using default '1. 000000‘ Loading weights from yolov 3. weights. . . Done! data/dog. jpg: Predicted in 0. 03 seconds. dog: 99% truck: 93% bicycle: 99% https: //pjreddie. com/darknet/yolo/
Comparison with Other Detectors
Recent Work on Yolo l Req-Yolo: Efficient quantization of the original network l l https: //arxiv. org/pdf/1909. 13396. pdf Yolo-Nano: Reducing size while increasing accuracy l https: //arxiv. org/pdf/1910. 01271 v 1. pdf [13] J. Redmon and A. Farhadi. YOLO 9000: better, faster, stronger. ar. Xiv preprint, 1612, 2016. [14] J. Redmon and A. Farhadi. Yolov 3: An incremental improvement. ar. Xiv preprint ar. Xiv: 1804. 02767, 2018.
Other CNNs Available in Darknet Model Darknet Tiny Darknet Ref Extraction Darknet 19 Darknet 53 Resnet 18 Resnet 34 Resnet 50 Resnet 101 Resnet 152 Top-1 Top-5 58. 7 61. 1 72. 5 72. 9 77. 2 70. 7 72. 4 75. 8 77. 1 77. 6 81. 7 83. 0 90. 8 91. 2 93. 8 89. 9 91. 1 92. 9 93. 7 93. 8 Tiny YOLOv 3 m. AP = 23. 7 YOLOv 3 -416 m. AP = 55. 3 Ops GPU 0. 98 B 0. 96 B 8. 52 B 7. 29 B 18. 57 B 4. 69 B 9. 52 B 9. 74 B 19. 70 B 29. 39 B 2. 9 ms 4. 8 ms 6. 2 ms 13. 7 ms 4. 6 ms 7. 1 ms 11. 4 ms 20. 0 ms 28. 6 ms Ops = 5. 41 B Ops = 65. 86 B CPU Weights 0. 14 s 0. 97 s 0. 87 s 2. 11 s 0. 57 s 1. 11 s 1. 13 s 2. 23 s 3. 31 s 4 MB 28 MB 90 MB 80 MB 159 MB 44 MB 83 MB 87 MB 160 MB 220 MB FPS = 244 FPS = 35
Overview l Introducing machine learning (ML) l l Implementation of ML l l Getting started, learning from the code Case studies l l l Data flow during training and evaluation The computation challenge Software and hardware implementation Darknet: A complete open-source ML solution l l Artificial neuron, neural network, convolution, etc Tiny Darknet: Image classification YOLO: Real-time object detection Further improving Darknet framework l l Speeding up computation Adding quantization
Further Improving Darknet Framework l It is possible to make Darknet much faster l Currently, matrix multiplication is used to compute convolution in both forward and backward passes l l l Efficient direct implementation of convolution can enable a 200 x speedup on a CPU, compared to the current CPU implementation l l PRO: Leads to fast implementation on GPU CON: Increases memory usage 9 x; layers need pre/post processing Making CPU implementation only ~5 x slower than GPU PRO: Can make training affordable for regular CPU users CON: Training may be not as fast as on GPU To achieve this speedup, update Darknet to use l Vectorized computation l l 8, 16, or 32 operations/cycle Multi-core implementation l Use cores to concurrently process different channels of the image
Quantization l Quantization turns floating-point computations into fixed-point ones: https: //sahnimanas. github. io/post/quantization-in-tflite/
Why Quantization? l Universality (broad applicability across models and use cases) l l Smaller model footprint l l Lower precision weights and activations allow for better cache reuse Faster computation l l With 8 -bit quantization, one can reduce the model size a factor of 4, with negligible accuracy loss. This leads to faster download times for models. Less working memory and cache for activations l l In many cases, one can start with an existing floating point model and quickly quantize it to obtain a fixed point quantized model with almost no accuracy loss, without needing to re-train the model. Many hardware platforms support fast inference with quantized values Most processors allow for faster processing of 8 -bit data Lower power l Moving 8 -bit data is 4 x more efficient than moving 32 -bit data. In many deep architectures, memory access dominates power consumption. https: //arxiv. org/pdf/1806. 08342. pdf
Implementation of Quantization l Quantization of floating-point number x [rmin; rmax] to k bits where l l l Nlevels = 2 k = (rmax - rmin) / (Nlevels – 1) De-quantization https: //arxiv. org/pdf/1806. 08342. pdf
Recent Work on Quantization l In 2018, the authors of https: //arxiv. org/abs/1805. 06085 stated l l In 2019, they say in "Accurate and efficient 2 -bit-quantized neural networks“ (https: //mlsys. org/Conferences/2019/doc/2019/168. pdf) l l “We show, for the first time, that both weights and activations can be quantized to 4 -bits of precision while still achieving accuracy comparable to full precision networks across a range of popular models and datasets”. “The combination of PACT [their 2018 paper] and SAWB results in a 2 bit QNN that achieves state-of-the-art classification accuracy (comparable to full precision networks) across a range of popular models and datasets”. Another recent paper (December 2019) refers to the above work, and introduces several orthogonal improvements, making results even better (https: //arxiv. org/abs/1912. 09356)
Res. Net 18 N | F 0 | F 1 | Type -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 | | | input | | | convo | | | maxpool | | | convo | 1 | | convo | | 2 | convo | | 4 | convo | | | convo | 6 | | convo | | 7 | convo | | 9 | convo | | | convo | 11 | | convo | | 12 | convo | | 14 | convo | | | convo | 16 | | convo | | 17 | convo | | 19 | convo | | | avepool | | | fullcon | | | softmax | | | 21 conv | Chan | K | G | P | R | Size | Addr | | 3 | | 64 | | 64 | | 128 | | 128 | | 256 | | 256 | | 512 | | 512 | | 1000 | | 5 | | 6448 | 0 7 3 1 3 3 3 3 7 1 0 | | | | | | | 0 0 0 0 0 0 0 | | | | | | | 0 3 0 0 1 1 1 1 0 0 0 | | | | | | | 0 1 0 0 1 1 1 1 0 1 0 | | | | | | | 0 2 2 1 1 1 2 2 1 1 1 0 | | | | | | | 224 112 56 56 56 28 28 28 14 14 14 7 7 7 1 1 1 | | | | | | | 1024 4096 1024 1024 512 512 512 256 256 256 128 128 128 250 1 Param | | | 9472 | | 4160 | 36928 | 8320 | 73856 | 147584 | 33024 | 295168 | 590080 | 131584 | 1180160 | 2359808 | |11688872 |44. 59 MB Activ | | | 802816 | | 200704 | 200704 | 100352 | 100352 | 50176 | 50176 | 25088 | 25088 | | 3037165 |11. 59 MB Macc | Cycle | | 118013952 | 460992 | | | 12845056 | 50176 | 115605504 | 451584 | 115605504 | 451584 | 6422528 | 25088 | 57802752 | 225792 | 115605504 | 451584 | 115605504 | 451584 | | | | 1826406400 | 7134400 | 3. 65 GOP |28. 0 FPS | | | | | | |
Darknet N | F 0 | F 1 | Type -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 | | | | | | | | | | | | | | | input convo maxpool convo maxpool convo avepool convo softmax | Chan | K | G | P | R | Size | Addr | | | | | 8 conv | 3 16 16 32 32 64 64 128 256 512 1024 1000 5 | | | | | 5072 | 0 3 2 3 2 3 2 3 4 1 0 | | | | | 0 0 0 0 0 | | | | | 0 1 1 1 1 0 0 0 | | | | | 0 2 0 2 0 0 0 | | | | | 0 1 2 1 2 1 2 1 0 | | | | | 256 128 64 64 32 32 16 16 8 8 4 4 1 1 1 Param | Activ | | 1024 | | | 448 | 1048576 | | 1024 | | | 4640 | 524288 | | 512 | | | 18496 | 262144 | | 256 | | | 73856 | 131072 | | 128 | | | 295168 | 65536 | | 64 | | | 1180160 | 32768 | | 32 | | 64 | 4719616 | 16384 | | 64 | | 62 | 1025000 | 1000 | | | | | 7317384 | 2795501 | | |27. 91 MB |10. 66 MB | Macc | Cycle | | | 28311552 | 110592 | | | 75497472 | 294912 | | | 75497472 | 294912 | | | 1024000 | | | 482320384 | 1884064| 0. 96 GOP |106. 2 FPS|
Tiny Darknet N | F 0 | F 1 | Type -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 | | | | | | | | | | | | | | | | | | | input convo maxpool convo convo convo avepool softmax | Chan | K | G | P | R | Size | Addr | | 3 | 0 | | 16 | 3 | | 16 | 2 | | 32 | 3 | | 32 | | 16 | 1 | | 128 | 3 | | 128 | 2 | | 32 | 1 | | 256 | 3 | | 256 | 2 | | 64 | 1 | | 512 | 3 | | 128 | 1 | | 1000 |14 | | 5 | 0 | | 16 conv | 4632 | | 0 0 0 0 0 0 | | | | | | | 0 1 1 0 1 0 0 | | | | | | | 0 2 0 2 2 2 0 0 0 | | | | | | | 0 1 2 1 1 1 0 0 | | | | | | | 224 112 56 56 56 28 28 28 14 14 1 1 | | | | | | | Param | Activ | 1024 | | | 448 | 802816 1024 | | | 4640 | 401408 512 | | 256 | 528 | 50176 2048 | 18560 | 401408 256 | 2064 | 50176 | 18560 | 401408 512 | | 128 | 4128 | 25088 1024 | 73984 | 200704 128 | 8224 | 25088 | 73984 | 200704 256 | | 64 | 16448 | 12544 512 | 295424 | 100352 64 | 32832 | 12544 512 | 295424 | 100352 128 | 65664 | 25088 1000 | 129000 | 196000 1000 | | 5 | | | 1039912 | 3608973 | 3. 97 MB |13. 77 MB | | | | | | | Macc | Cycle | | | 21676032 | 84672 | | | 57802752 | 225792 | | | 1605632 | 6272 | 57802752 | 225792 | 6422528 | 25088 | 57802752 | 225792 | | | 3211264 | 12544 | 57802752 | 225792 | 6422528 | 25088 | 57802752 | 225792 | 12845056 | 50176 | 25088000 | 98000 | | | 491524096 | 1920016 | 0. 98 GOP |104. 2 FPS|
Tiny Yolo. V 3 N | F 0 | F 1 | Type -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 | | | | | | | | | | 13 | | | | | | | | | 8 | | | input convo maxpool convo maxpool convo yolo convo interp concat convo yolo | Chan | K | G | P | R | Size | Addr | | 3 | 0 | 0 | 0 | | 16 | 3 | 0 | 1 | 2 | 1 | | 16 | 2 | 0 | 1 | 0 | 2 | | 32 | 3 | 0 | 1 | 2 | 1 | | 32 | 0 | 1 | 0 | 2 | | 64 | 3 | 0 | 1 | 2 | 1 | | 64 | 2 | 0 | 1 | 0 | 2 | | 128 | 3 | 0 | 1 | 2 | 1 | | 128 | 2 | 0 | 1 | 0 | 2 | | 256 | 3 | 0 | 1 | 2 | 1 | | 256 | 2 | 0 | 1 | 0 | 2 | | 512 | 3 | 0 | 1 | 2 | 1 | | 512 | 0 | 1 | | 1024 | 3 | 0 | 1 | 2 | 1 | | 256 | 1 | 0 | 2 | 1 | | 512 | 3 | 0 | 1 | 2 | 1 | | 255 | 1 | 0 | 0 | 1 | | 255 | 0 | 0 | 0 | | 128 | 1 | 0 | 2 | 1 | | 128 | 0 | 0 | 2 | | 384 | 0 | 0 | 0 | | 256 | 3 | 0 | 1 | 2 | 1 | | 255 | 1 | 0 | 0 | 1 | | 255 | 0 | 0 | 0 | | | | 13 conv | 5727 | | | 416 208 104 52 52 26 26 13 13 13 26 26 26 | | | | | | | Param | 4096 | | 448 4096 | | 4640 2048 | | 18496 1024 | | 73856 512 | | 295168 256 | | 1180160 512 | 1024 | 4719616 256 | 262400 512 | 1180160 255 | 130815 255 | 128 | 32896 512 | 1536 | 1024 | 884992 1020 | 65535 1020 | | 8849182 |33. 76 MB Activ | | | 2768896 | | 1384448 | | 692224 | | 346112 | | 173056 | | 86528 | | 173056 | 43264 | 86528 | 43095 | | 21632 | | | 173056 | 172380 | | 8457267 |32. 26 MB Macc | Cycle | | 74760192 | 292032 | | | 199360512 | 778752 | | | 199360512 | 778752 | | | 797442048 | 3115008 | 44302336 | 173056 | 199360512 | 778752 | 22064640 | 86190 | | | 5537792 | 21632 | | | 598081536 | 2336256 | 44129280 | 172380 | | | 2782480896 |10869066 | 5. 56 GOP |18. 4 FPS | | | | | | |
Yolo. V 3 N | F 0 | F 1 | Type -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 | | | | | | | | | | | | | | | | | | | | 1 4 6 9 11 13 15 17 19 21 23 26 28 30 32 34 | | | | | | | | | | input convo convo convo convo convo convo convo convo convo convo | Chan | K | G | P | R | Size | Addr | | | | | | | | | | 3 32 64 128 256 128 256 128 256 512 256 512 256 | | | | | | | | | | 0 3 3 1 3 1 3 1 3 3 1 3 1 3 1 | | | | | | | | | | 0 0 0 0 0 0 0 0 0 0 | | | | | | | | | | 0 1 1 0 1 0 1 1 0 1 0 1 0 | | | | | | | | | | 0 2 2 2 2 2 2 2 2 2 2 | | | | | | | | | | 0 1 2 1 1 1 1 1 1 1 1 2 1 1 1 | | | | | | | | | | 608 304 304 152 152 152 76 76 76 76 76 38 38 38 | 4096 |32768 |16384 | 8192 | 4096 | 2048 | 4096 | 2048 | 4096 | 2048 | 1024 | 2048 | 1024 | | | | | | | | | | Param | 896 18496 2080 18496 73856 8256 73856 295168 32896 295168 32896 295168 1180160 131328 1180160 131328 Activ | | |11829248 | 5914624 | 2957312 | 1478656 | 739328 | 1478656 | 739328 | 1478656 | 739328 | 369664 | 739328 | 369664 | | | | | | | | | | Macc | 319389696 1703411712 189267968 1703411712 189267968 1703411712 189267968 1703411712 189267968 1703411712 189267968 | | | | | | | | | | Cycle | 1247616 6653952 739328 6653952 739328 6653952 739328 6653952 739328 6653952 739328 | | | | | | | | | |
N | F 0 | F 1 | Type 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 | | 36 | convo | | 38 | convo | | 40 | convo | | 43 | convo | | 45 | convo | | 47 | convo | | 49 | convo | | | convo | | | yolo | 56 | | convo | | | interp | | 42 | concat | | | convo | | | convo | | | yolo | 67 | | convo | | | interp | | 25 | concat | | | convo | | | convo | | | yolo | | | 75 conv | Chan | K | G | P | R | Size | Addr | | 512 | 256 | 512 | 1024 | 512 | 1024 | 255 | 256 | 768 | 256 | 512 | 255 | 128 | 384 | 128 | 256 | 255 | |29373 | | | | | | | | | | | | 3 1 3 1 3 1 0 1 0 0 1 3 1 3 1 3 1 0 | | | | | | | | | | | | 0 0 0 0 0 0 0 0 0 0 0 | | | | | | | | | | | | 1 0 1 0 1 0 0 0 0 0 0 1 0 1 0 1 0 0 | | | | | | | | | | | | 2 2 2 2 2 0 0 2 2 2 2 2 2 0 0 | | | | | | | | | | | | 1 1 1 2 1 1 1 1 0 1 2 0 1 1 1 1 0 | | | | | | | | | | | | 38 38 38 19 19 19 19 19 38 38 38 76 76 76 | | | | | | | | | | | | 2048 1024 512 1024 512 1024 255 256 1024 3072 1024 2048 1020 512 2048 6144 2048 4096 4080 Param | Activ | Macc | Cycle | | 1180160 | 739328 | 1703411712 | 6653952 | | 131328 | 369664 | 189267968 | 739328 | | 1180160 | 739328 | 1703411712 | 6653952 | | 4719616 | 369664 | 1703411712 | 6653952 | | 524800 | 184832 | 189267968 | 739328 | | 4719616 | 369664 | 1703411712 | 6653952 | | 524800 | 184832 | 189267968 | 739328 | | 4719616 | 369664 | 1703411712 | 6653952 | | 261375 | 92055 | 94264320 | 368220 | | | | 131328 | 92416 | 47316992 | 184832 | | | 196864 | 369664 | 283901952 | 1108992 | | 1180160 | 739328 | 1703411712 | 6653952 | | 131328 | 369664 | 189267968 | 739328 | | 1180160 | 739328 | 1703411712 | 6653952 | | 130815 | 368220 | 188528640 | 736440 | | | | 32896 | 184832 | 47316992 | 184832 | | | 49280 | 739328 | 283901952 | 1108992 | | 295168 | 1478656 | 1703411712 | 6653952 | | 32896 | 739328 | 189267968 | 739328 | | 295168 | 1478656 | 1703411712 | 6653952 | | 65535 | 1472880 | 377057280 | 1472880 | | | |61922845 |89266275 |70345950208 |274788868| |236. 22 MB|340. 52 MB| 140. 69 GOP| 0. 7 FPS|
- Slides: 45