Details of the Researcher

PHOTO

Hiroki Nakahara
Section
Unprecedented-scale Data Analytics Center
Job title
Professor
e-Rad No.
20624414

Research History 5

  • 2023/10 - Present
    Tohoku University Unprecedented-scale Data Analytics Center Professor

  • 2016/04 - 2023/09
    Tokyo Institute of Technology Associate Professor

  • 2014/10 - 2016/03
    Ehime University Senior Assistant Professor

  • 2012/10 - 2014/09
    Kagoshima University Assistant Professor

  • 2007/09 - 2012/09
    Kyushu Institute of Technology Faculty of Computer Science and Systems Engineering, Department of Computer Science and Electronics

Professional Memberships 3

  • IPSJ

  • IEICE

  • IEEE

Research Interests 7

  • Machine Learning

  • Multiple Valued Logic

  • Deep Learning

  • Computer System

  • Embedded System

  • Reconfigurable System

  • FPGA

Research Areas 4

  • Manufacturing technology (mechanical, electrical/electronic, chemical engineering) / Communication and network engineering /

  • Informatics / Information networks /

  • Informatics / Software /

  • Informatics / Information theory /

Awards 12

  1. 最優秀エンジニア講演賞

    2018 Design Solution Forum 2018

  2. 最優秀エンジニア講演賞

    2017 Design Solution Forum 2017

  3. Best Demo Award

    2017 IEEE/ACM International Workshop on Reconfigurable Architecture

  4. 最優秀論文賞

    2016 多値論理フォーラム

  5. Young Researcher Award

    2015 IEEE CASS Shikoku Chapter

  6. Kenneth C. Smith Early Career Award

    2014 IEEE International Symposium on ISMVL2013

  7. Best Paper Award

    2013 IEEE 7th International Symposium on MCSoC-13

  8. Funai Best Paper Award

    2012 FIT2012

  9. MEMOCODE2010 Design Contest Winner Award

    2010

  10. SASIMI2010 Outstanding Paper Award

    2010

  11. デザインガイア2009最優秀ポスター発表賞

    2009

  12. Excellent Student Award of the IEEE Fukuoka Section Award

    2006

Show all ︎Show 5

Papers 78

  1. A Novel Data Representation Towards Efficient FPGA-based Quantum Computer Simulation.

    Haruhiko Hasegawa, Masayuki Shimoda, Hiroki Nakahara, Takefumi Miyoshi

    ISMVL 117-122 2025

    DOI: 10.1109/ISMVL64713.2025.00031  

  2. Edge Inference Engine for Deep & Random Sparse Neural Networks with 4-bit Cartesian-Product MAC Array and Pipelined Activation Aligner

    Kota Ando, Jaehoon Yu, Kazutoshi Hirose, Hiroki Nakahara, Kazushi Kawamura, Thiem Van Chu, Masato Motomura

    2021 IEEE Hot Chips 33 Symposium (HCS) 2021/08/22

    Publisher: IEEE

    DOI: 10.1109/hcs52781.2021.9567328  

  3. FPGA-Based Inter-layer Pipelined Accelerators for Filter-Wise Weight-Balanced Sparse Fully Convolutional Networks with Overlapped Tiling.

    Masayuki Shimoda, Youki Sada, Hiroki Nakahara

    Journal of Signal Processing Systems 93 (5) 499-512 2021

    DOI: 10.1007/s11265-021-01642-6  

  4. Fast Monocular Depth Estimation on an FPGA.

    Youki Sada, Naoto Soga, Masayuki Shimoda, Akira Jinguji, Shimpei Sato, Hiroki Nakahara

    2020 IEEE International Parallel and Distributed Processing Symposium Workshops 143-146 2020

    Publisher: IEEE

    DOI: 10.1109/IPDPSW50202.2020.00032  

  5. SENTEI: Filter-Wise Pruning with Distillation towards Efficient Sparse Convolutional Neural Network Accelerators.

    Masayuki Shimoda, Youki Sada, Ryosuke Kuramochi, Shimpei Sato, Hiroki Nakahara

    IEICE Transactions on Information & Systems 103-D (12) 2463-2470 2020

    DOI: 10.1587/transinf.2020PAP0013  

  6. Guinness: A GUI based binarized deep neural network framework for software programmers

    Hiroki Nakahara, Haruyoshi Yonekawa, Tomoya Fujii, Masayuki Shimoda, Shimpei Sato

    IEICE Transactions on Information and Systems E102D (5) 1003-1011 2019/05/01

    Publisher: Institute of Electronics, Information and Communication, Engineers, IEICE

    DOI: 10.1587/transinf.2018RCP0002  

    ISSN: 1745-1361 0916-8532

  7. Many Universal Convolution Cores for Ensemble Sparse Convolutional Neural Networks.

    Ryosuke Kuramochi, Youki Sada, Masayuki Shimoda, Shimpei Sato, Hiroki Nakahara

    13th IEEE International Symposium on Embedded Multicore/Many-core Systems-on-Chip(MCSoC) 93-100 2019

    Publisher: IEEE

    DOI: 10.1109/MCSoC.2019.00021  

  8. A Dataflow Pipelining Architecture for Tile Segmentation with a Sparse MobileNet on an FPGA.

    Youki Sada, Masayuki Shimoda, Akira Jinguji, Hiroki Nakahara

    International Conference on Field-Programmable Technology(FPT) 267-270 2019

    Publisher: IEEE

    DOI: 10.1109/ICFPT47387.2019.00044  

  9. An FPGA Implementation of Real-Time Object Detection with a Thermal Camera.

    Masayuki Shimoda, Youki Sada, Ryosuke Kuramochi, Hiroki Nakahara

    29th International Conference on Field Programmable Logic and Applications(FPL) 413-414 2019

    Publisher: IEEE

    DOI: 10.1109/FPL.2019.00072  

  10. FPGA-Based Training Accelerator Utilizing Sparseness of Convolutional Neural Network.

    Hiroki Nakahara, Youki Sada, Masayuki Shimoda, Kouki Sayama, Akira Jinguji, Shimpei Sato

    29th International Conference on Field Programmable Logic and Applications(FPL) 180-186 2019

    Publisher: IEEE

    DOI: 10.1109/FPL.2019.00036  

  11. An FPGA-based Fine Tuning Accelerator for a Sparse CNN.

    Hiroki Nakahara, Akira Jinguji, Masayuki Shimoda, Shimpei Sato

    Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays(FPGA) 186-186 2019

    Publisher: ACM

    DOI: 10.1145/3289602.3293967  

  12. Filter-Wise Pruning Approach to FPGA Implementation of Fully Convolutional Network for Semantic Segmentation.

    Masayuki Shimoda, Youki Sada, Hiroki Nakahara

    Applied Reconfigurable Computing - 15th International Symposium(ARC) 371-386 2019

    Publisher: Springer

    DOI: 10.1007/978-3-030-17227-5_26  

  13. Power Efficient Object Detector with an Event-Driven Camera for Moving Object Surveillance on an FPGA.

    Masayuki Shimoda, Shimpei Sato, Hiroki Nakahara

    IEICE Transactions on Information & Systems 102-D (5) 1020-1028 2019

    DOI: 10.1587/transinf.2018RCP0005  

  14. A Tri-State Weight Convolutional Neural Network for an FPGA: Applied to YOLOv2 Object Detector

    Hiroki Nakahara, Masayuki Shimoda, Shimpei Sato

    2018 International Conference on Field-Programmable Technology (FPT) 298-301 2018/12

    Publisher: IEEE

    DOI: 10.1109/fpt.2018.00058  

  15. BRein Memory: A Single-Chip Binary/Ternary Reconfigurable in-Memory Deep Neural Network Accelerator Achieving 1.4 TOPS at 0.6 W Peer-reviewed

    Kota Ando, Kodai Ueyoshi, Kentaro Orimo, Haruyoshi Yonekawa, Shimpei Sato, Hiroki Nakahara, Shinya Takamaeda-Yamazaki, Masayuki Ikebe, Tetsuya Asai, Tadahiro Kuroda, Masato Motomura

    IEEE Journal of Solid-State Circuits 53 (4) 983-994 2018/04/01

    Publisher: Institute of Electrical and Electronics Engineers Inc.

    DOI: 10.1109/JSSC.2017.2778702  

    ISSN: 0018-9200

  16. A threshold neuron pruning for a binarized deep neural network on an FPGA Peer-reviewed

    Tomoya Fujii, Shimpei Sato, Hiroki Nakahara

    IEICE Transactions on Information and Systems E101D (2) 376-386 2018/02/01

    Publisher: Institute of Electronics, Information and Communication, Engineers, IEICE

    DOI: 10.1587/transinf.2017RCP0013  

    ISSN: 1745-1361 0916-8532

  17. An FPGA realization of a random forest with k-means clustering using a high-level synthesis design Peer-reviewed

    Akira Jinguji, Shimpei Sato, Hiroki Nakahara

    IEICE Transactions on Information and Systems E101D (2) 354-362 2018/02/01

    Publisher: Institute of Electronics, Information and Communication, Engineers, IEICE

    DOI: 10.1587/transinf.2017RCP0006  

    ISSN: 1745-1361 0916-8532

  18. New Generation Dynamically Reconfigurable Processor Technology for Accelerating Embedded AI Applications. Peer-reviewed

    Taro Fujii, Takao Toi, Teruhito Tanaka, Katsumi Togawa, Toshiro Kitaoka, Kengo Nishino, Noritsugu Nakamura, Hiroki Nakahara, Masato Motomura

    2018 IEEE Symposium on VLSI Circuits, Honolulu, HI, USA, June 18-22, 2018 41-42 2018

    Publisher: IEEE

    DOI: 10.1109/VLSIC.2018.8502438  

  19. A Ternary Weight Binary Input Convolutional Neural Network: Realization on the Embedded Processor. Peer-reviewed

    Haruyoshi Yonekawa, Shimpei Sato, Hiroki Nakahara

    48th IEEE International Symposium on Multiple-Valued Logic, ISMVL 2018, Linz, Austria, May 16-18, 2018 174-179 2018

    Publisher: IEEE

    DOI: 10.1109/ISMVL.2018.00038  

  20. A High-speed Low-power Deep Neural Network on an FPGA based on the Nested RNS: Applied to an Object Detector. Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao

    IEEE International Symposium on Circuits and Systems, ISCAS 2018, 27-30 May 2018, Florence, Italy 1-5 2018

    Publisher: IEEE

    DOI: 10.1109/ISCAS.2018.8351850  

  21. A Performance Per Power Efficient Object Detector on an FPGA for Robot Operating System (ROS). Peer-reviewed

    Haoxuan Cheng, Shimpei Sato, Hiroki Nakahara

    Proceedings of the 9th International Symposium on Highly-Efficient Accelerators and Reconfigurable Technologies, HEART 2018, Toronto, ON, Canada, June 20-22, 2018 20:1-20:4 2018

    Publisher: ACM

    DOI: 10.1145/3241793.3241814  

  22. Power Efficient Object Detector with an Event-Driven Camera on an FPGA. Peer-reviewed

    Masayuki Shimoda, Shimpei Sato, Hiroki Nakahara

    Proceedings of the 9th International Symposium on Highly-Efficient Accelerators and Reconfigurable Technologies, HEART 2018, Toronto, ON, Canada, June 20-22, 2018 10:1-10:6 2018

    Publisher: ACM

    DOI: 10.1145/3241793.3241803  

  23. Demonstration of Object Detection for Event-Driven Cameras on FPGAs and GPUs. Peer-reviewed

    Masayuki Shimoda, Shimpei Sato, Hiroki Nakahara

    28th International Conference on Field Programmable Logic and Applications, FPL 2018, Dublin, Ireland, August 27-31, 2018 461-462 2018

    Publisher: IEEE

    DOI: 10.1109/FPL.2018.00090  

  24. A Demonstration of FPGA-Based You Only Look Once Version2 (YOLOv2). Peer-reviewed

    Hiroki Nakahara, Masayuki Shimoda, Shimpei Sato

    28th International Conference on Field Programmable Logic and Applications, FPL 2018, Dublin, Ireland, August 27-31, 2018 457-458 2018

    Publisher: IEEE

    DOI: 10.1109/FPL.2018.00088  

  25. A Lightweight YOLOv2: A Binarized CNN with A Parallel Support Vector Regression for an FPGA. Peer-reviewed

    Hiroki Nakahara, Haruyoshi Yonekawa, Tomoya Fujii, Shimpei Sato

    Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, FPGA 2018, Monterey, CA, USA, February 25-27, 2018 31-40 2018

    Publisher: ACM

    DOI: 10.1145/3174243.3174266  

  26. A fully connected layer elimination for a binarizec convolutional neural network on an FPGA Peer-reviewed

    Hiroki Nakahara, Tomoya Fujii, Shimpei Sato

    2017 27th International Conference on Field Programmable Logic and Applications, FPL 2017 1-4 2017/10/02

    Publisher: Institute of Electrical and Electronics Engineers Inc.

    DOI: 10.23919/FPL.2017.8056771  

  27. A Random Forest Using a Multi-valued Decision Diagram on an FPGA Peer-reviewed

    Hiroki Nakahara, Akira Jinguji, Simpei Sato, Tsutomu Sasao

    Proceedings of The International Symposium on Multiple-Valued Logic 266-271 2017/06/30

    Publisher: IEEE Computer Society

    DOI: 10.1109/ISMVL.2017.40  

    ISSN: 0195-623X

  28. On-chip memory based binarized convolutional deep neural network applying batch normalization free technique on an FPGA Peer-reviewed

    Haruyoshi Yonekawa, Hiroki Nakahara

    Proceedings - 2017 IEEE 31st International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2017 98-105 2017/06/30

    Publisher: Institute of Electrical and Electronics Engineers Inc.

    DOI: 10.1109/IPDPSW.2017.95  

  29. In-memory area-efficient signal streaming processor design for binary neural networks.

    Haruyoshi Yonekawa, Shimpei Sato, Hiroki Nakahara, Kota Ando, Kodai Ueyoshi, Kazutoshi Hirose, Kentaro Orimo, Shinya Takamaeda-Yamazaki, Masayuki Ikebe, Tetsuya Asai, Masato Motomura

    IEEE 60th International Midwest Symposium on Circuits and Systems(MWSCAS) 116-119 2017

    Publisher: IEEE

    DOI: 10.1109/MWSCAS.2017.8052874  

  30. All binarized convolutional neural network and its implementation on an FPGA. Peer-reviewed

    Masayuki Shimoda, Shimpei Sato, Hiroki Nakahara

    International Conference on Field Programmable Technology, ICFPT 2017, Melbourne, Australia, December 11-13, 2017 291-294 2017

    Publisher: IEEE

    DOI: 10.1109/FPT.2017.8280163  

  31. An object detector based on multiscale sliding window search using a fully pipelined binarized CNN on an FPGA. Peer-reviewed

    Hiroki Nakahara, Haruyoshi Yonekawa, Shimpei Sato

    International Conference on Field Programmable Technology, ICFPT 2017, Melbourne, Australia, December 11-13, 2017 168-175 2017

    Publisher: IEEE

    DOI: 10.1109/FPT.2017.8280135  

  32. A Batch Normalization Free Binarized Convolutional Deep Neural Network on an FPGA (Abstract Only). Peer-reviewed

    Hiroki Nakahara, Haruyoshi Yonekawa, Hisashi Iwamoto, Masato Motomura

    Proceedings of the 2017 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, FPGA 2017, Monterey, CA, USA, February 22-24, 2017 290-290 2017

    Publisher: ACM

  33. An FPGA realization of a deep convolutional neural network using a threshold neuron pruning Peer-reviewed

    Tomoya Fujii, Simpei Sato, Hiroki Nakahara, Masato Motomura

    Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 10216 268-280 2017

    Publisher: Springer Verlag

    DOI: 10.1007/978-3-319-56258-2_23  

    ISSN: 1611-3349 0302-9743

  34. LUT Cascades Based on Edge-Valued Multi-Valued Decision Diagrams: Application to Packet Classification Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Hisashi Iwamoto, Munehiro Matsuura

    IEEE JOURNAL ON EMERGING AND SELECTED TOPICS IN CIRCUITS AND SYSTEMS 6 (1) 73-86 2016/03

    DOI: 10.1109/JETCAS.2016.2528638  

    ISSN: 2156-3357

  35. An FFT Circuit Using Nested RNS in a Digital Spectrometer for a Radio Telescope Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Hiroyuki Nakanishi, Kazumasa Iwai, Tohru Nagao, Naoya Ogawa

    2016 IEEE 46TH INTERNATIONAL SYMPOSIUM ON MULTIPLE-VALUED LOGIC (ISMVL 2016) 60-65 2016

    DOI: 10.1109/ISMVL.2016.35  

    ISSN: 0195-623X

  36. An Acceleration of a Random Forest Classification using Altera SDK for OpenCL Peer-reviewed

    Hiroki Nakahara, Akira Jinguji, Tomonori Fujii, Simpei Sato

    2016 INTERNATIONAL CONFERENCE ON FIELD-PROGRAMMABLE TECHNOLOGY (FPT) 289-292 2016

    DOI: 10.1109/FPT.2016.7929555  

  37. A Memory-Based Realization of a Binarized Deep Convolutional Neural Network Peer-reviewed

    Hiroki Nakahara, Haruyoshi Yonekawa, Tsutomu Sasao, Hisashi Iwamoto, Masato Motomura

    2016 INTERNATIONAL CONFERENCE ON FIELD-PROGRAMMABLE TECHNOLOGY (FPT) 277-280 2016

    DOI: 10.1109/FPT.2016.7929552  

  38. An FFT Circuit for a Spectrometer of a Radio Telescope using the Nested RNS including the Constant Division. Peer-reviewed

    Hiroki Nakahara, Hiroyuki Nakanishi, Kazumasa Iwai, Tsutomu Sasao

    SIGARCH Computer Architecture News 44 (4) 44-49 2016

    DOI: 10.1145/3039902.3039911  

  39. An Update Method for a Low Power CAM Emulator Using an LUT Cascade Based on an EVMDD (k) Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura, Hisashi Iwamoto

    JOURNAL OF MULTIPLE-VALUED LOGIC AND SOFT COMPUTING 26 (1-2) 109-123 2016

    ISSN: 1542-3980

    eISSN: 1542-3999

  40. An RNS FFT Circuit Using LUT Cascades Based on a Modulo EVMDD Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Hiroyuki Nakanishi, Kazumasa Iwai

    Proceedings of The International Symposium on Multiple-Valued Logic 2015- 97-102 2015/09/02

    Publisher: IEEE Computer Society

    DOI: 10.1109/ISMVL.2015.41  

    ISSN: 0195-623X

  41. A Memory-Based IPv6 Lookup Architecture Using Parallel Index Generation Units Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura, Hisashi Iwamoto, Yasuhiro Terao

    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS E98D (2) 262-271 2015/02

    DOI: 10.1587/transinf.2014RCP0006  

    ISSN: 1745-1361

  42. A Deep Convolutional Neural Network Based on Nested Residue Number System Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao

    2015 25TH INTERNATIONAL CONFERENCE ON FIELD PROGRAMMABLE LOGIC AND APPLICATIONS 1-6 2015

    DOI: 10.1109/FPL.2015.7293933  

    ISSN: 1946-1488

  43. A dynamically reconfigurable mixed analog-digital filter bank Peer-reviewed

    Hiroki Nakahara, Hideki Yoshida, Shin-Ich Shioya, Renji Mikami, Tsutomu Sasao

    Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 9040 267-279 2015

    Publisher: Springer Verlag

    DOI: 10.1007/978-3-319-16214-0_22  

    ISSN: 1611-3349 0302-9743

  44. A Packet Classifier Based on Prefetching EVMDD (k) Machines Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS E97D (9) 2243-2252 2014/09

    DOI: 10.1587/transinf.2013LOP0020  

    ISSN: 1745-1361

  45. An AWF Digital Spectrometer for a Radio Telescope Peer-reviewed

    Hiroki Nakahara, Hiroyuki Nakanishi, Kazumasa Iwai

    2014 INTERNATIONAL CONFERENCE ON RECONFIGURABLE COMPUTING AND FPGAS (RECONFIG) 1-6 2014

    DOI: 10.1109/ReConFig.2014.7032503  

    ISSN: 2325-6532

  46. AN UPDATE METHOD FOR A CAM EMULATOR USING AN LUT CASCADE BASED ON AN EVMDD (K) Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    2014 IEEE 44TH INTERNATIONAL SYMPOSIUM ON MULTIPLE-VALUED LOGIC (ISMVL 2014) 1-6 2014

    DOI: 10.1109/ISMVL.2014.9  

    ISSN: 0195-623X

  47. Automatic Adjustment System for Optical Interconnection Transmitter Using Improved Particle Swarm Optimization Peer-reviewed

    Kenichi Ohhata, Hiroki Nakahara, Takuya Inoue, Toru Yazaki, Norio Chujo, Takuma Nishimoto

    2014 14TH INTERNATIONAL SYMPOSIUM ON INTEGRATED CIRCUITS (ISIC) 584-587 2014

    DOI: 10.1109/ISICIR.2014.7029441  

    ISSN: 2325-0631

  48. A Heterogeneous Multi-valued Decision Diagram Machine for Encoded Characteristic Function for Non-zero Outputs Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    JOURNAL OF MULTIPLE-VALUED LOGIC AND SOFT COMPUTING 23 (3-4) 365-377 2014

    ISSN: 1542-3980

  49. A packet classifier using parallel EVMDD (k) machine Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    Proceedings - IEEE 7th International Symposium on Embedded Multicore/Manycore System-on-Chip, MCSoC 2013 43-48 2013

    Publisher: IEEE Computer Society

    DOI: 10.1109/MCSoC.2013.26  

  50. A machine to evaluate decomposed multi-terminal multi-valued decision diagrams for characteristic functions Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    Proceedings of The International Symposium on Multiple-Valued Logic 90-95 2013

    Publisher: IEEE Computer Society

    DOI: 10.1109/ISMVL.2013.6  

    ISSN: 0195-623X

  51. A HIGH-SPEED FFT BASED ON A SIX-STEP ALGORITHM: APPLIED TO A RADIO TELESCOPE FOR A SOLAR RADIO BURST Peer-reviewed

    Hiroki Nakahara, Kazumasa Iwai, Hiroyuki Nakanishi

    PROCEEDINGS OF THE 2013 INTERNATIONAL CONFERENCE ON FIELD-PROGRAMMABLE TECHNOLOGY (FPT) 430-433 2013

    DOI: 10.1109/FPT.2013.6718406  

  52. A PACKET CLASSIFIER USING LUT CASCADES BASED ON EVMDDS (K) Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    2013 23RD INTERNATIONAL CONFERENCE ON FIELD PROGRAMMABLE LOGIC AND APPLICATIONS (FPL 2013) PROCEEDINGS 1-6 2013

    DOI: 10.1109/FPL.2013.6645518  

    ISSN: 1946-1488

  53. An architecture for IPv6 lookup using parallel index generation units Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 7806 59-71 2013

    Publisher: Springer

    DOI: 10.1007/978-3-642-36812-7_6  

    ISSN: 0302-9743 1611-3349

  54. A virus scanning engine using an MPU and an IGU based on row-shift decomposition Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    IEICE Transactions on Information and Systems E96-D (8) 1667-1675 2013

    Publisher: Institute of Electronics, Information and Communication, Engineers, IEICE

    DOI: 10.1587/transinf.E96.D.1667  

    ISSN: 1745-1361 0916-8532

  55. A regular expression matching circuit: Decomposed non-deterministic realization with prefix sharing and multi-character transition Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    MICROPROCESSORS AND MICROSYSTEMS 36 (8) 644-664 2012/11

    DOI: 10.1016/j.micpro.2012.05.009  

    ISSN: 0141-9331

    eISSN: 1872-9436

  56. A Design Method of a Regular Expression Matching Circuit Based on Decomposed Automaton Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS E95D (2) 364-373 2012/02

    DOI: 10.1587/transinf.E95.D.364  

    ISSN: 1745-1361

  57. Multi-Terminal Multi-Valued Decision Diagrams for Characteristic Function Representing Cluster Decomposition Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    2012 42ND IEEE INTERNATIONAL SYMPOSIUM ON MULTIPLE-VALUED LOGIC (ISMVL) 148-153 2012

    DOI: 10.1109/ISMVL.2012.45  

    ISSN: 0195-623X

  58. On a wideband fast fourier transform using piecewise linear approximations: Application to a radio telescope spectrometer Peer-reviewed

    Hiroki Nakahara, Hiroyuki Nakanishi, Tsutomu Sasao

    Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 7439 (1) 202-217 2012

    Publisher: Springer

    DOI: 10.1007/978-3-642-33078-0_15  

    ISSN: 0302-9743 1611-3349

    eISSN: 1611-3349

  59. A low-cost and high-performance virus scanning engine using a binary CAM emulator and an MPU Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 7199 202-214 2012

    Publisher: Springer

    DOI: 10.1007/978-3-642-28365-9_17  

    ISSN: 0302-9743 1611-3349

  60. On a wideband fast fourier transform for a radio telescope. Peer-reviewed

    Hiroki Nakahara, Hiroyuki Nakanishi, Tsutomu Sasao

    SIGARCH Computer Architecture News 40 (5) 46-51 2012

    DOI: 10.1145/2460216.2460225  

  61. A Comparison of Multi-Valued and Heterogeneous Decision Diagram Machines Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    JOURNAL OF MULTIPLE-VALUED LOGIC AND SOFT COMPUTING 19 (1-3) 203-217 2012

    ISSN: 1542-3980

  62. On a Prefetching Heterogeneous MDD Machine Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    2011 IEEE 54TH INTERNATIONAL MIDWEST SYMPOSIUM ON CIRCUITS AND SYSTEMS (MWSCAS) 2011

    ISSN: 1548-3746

  63. A Comparison of Heterogeneous Multi-valued Decision Diagram Machines for Multiple-output Logic Functions Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    2011 41ST IEEE INTERNATIONAL SYMPOSIUM ON MULTIPLE-VALUED LOGIC (ISMVL) 125-130 2011

    DOI: 10.1109/ISMVL.2011.15  

    ISSN: 0195-623X

  64. A Regular Expression Matching Circuit Based on a Decomposed Automaton Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    RECONFIGURABLE COMPUTING: ARCHITECTURES, TOOLS AND APPLICATIONS 6578 16-28 2011

    DOI: 10.1007/978-3-642-19475-7_4  

    ISSN: 0302-9743

  65. A regular expression matching circuit based on a modular non-deterministic finite automaton with multi-character transition

    H. Nakahara, M. Matsuura

    The 16th Workshop on Synthesis And System Integration of Mixed Information technologies (SASIMI-2010) 359 - 364-364 2010/10

  66. A regular expression matching using non-deterministic finite automaton

    Hiroki Nakahara

    Proc. of Eighth ACM/IEEE International Conference on Formal Methods and Models for Codesign, (MEMOCODE-2010), (IEEE, USA) 73-76 2010/07

    DOI: 10.1109/MEMCOD.2010.5558621  

  67. A realization of index generation functions using modules of uniform sizes

    M. Matsuura, H. Nakahara

    19th International Workshop on Logic and Synthesis (IWLS-2010) 201-208 2010/06

  68. A Comparison of Architectures for Various Decision Diagram Machines Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    40TH IEEE INTERNATIONAL SYMPOSIUM ON MULTIPLE-VALUED LOGIC ISMVL 2010 229-234 2010

    DOI: 10.1109/ISMVL.2010.50  

    ISSN: 0195-623X

  69. A Packet Classifier Using a Parallel Branching Program Machine Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    13TH EUROMICRO CONFERENCE ON DIGITAL SYSTEM DESIGN: ARCHITECTURES, METHODS AND TOOLS 745-752 2010

    DOI: 10.1109/DSD.2010.18  

  70. A Quaternary Decision Diagram Machine and the Optimization of Its Code Peer-reviewed

    Tsutomu Sasao, Hiroki Nakahara, Munehiro Matsuura, Yoshifumi Kawamura, Jon T. Butler

    ISMVL: 2009 39TH IEEE INTERNATIONAL SYMPOSIUM ON MULTIPLE-VALUED LOGIC 362-+ 2009

    DOI: 10.1109/ISMVL.2009.35  

  71. A Virus Scanning Engine Using a Parallel Finite-Input Memory Machine and MPUs Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura, Yoshifumi Kawamura

    FPL: 2009 INTERNATIONAL CONFERENCE ON FIELD PROGRAMMABLE LOGIC AND APPLICATIONS 635-+ 2009

    DOI: 10.1109/FPL.2009.5272396  

    ISSN: 1946-1488

  72. The Parallel Sieve Method for a Virus Scanning Engine Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura, Yoshifumi Kawamura

    PROCEEDINGS OF THE 2009 12TH EUROMICRO CONFERENCE ON DIGITAL SYSTEM DESIGN, ARCHITECTURES, METHODS AND TOOLS 809-+ 2009

    DOI: 10.1109/DSD.2009.208  

  73. A Parallel Branching Program Machine for Emulation of Sequential Circuits Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura, Yoshifumi Kawamura

    RECONFIGURABLE COMPUTING: ARCHITECTURES, TOOLS AND APPLICATIONS 5453 261-+ 2009

    DOI: 10.1007/978-3-642-00641-8_26  

    ISSN: 0302-9743

  74. A CAM emulator using look-up table cascades Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    Proceedings - 21st International Parallel and Distributed Processing Symposium, IPDPS 2007; Abstracts and CD-ROM 1-8 2007

    Publisher: IEEE

    DOI: 10.1109/IPDPS.2007.370372  

  75. Implementations of reconfigurable logic arrays on FPGAs Peer-reviewed

    Tsutomu Sasao, Hiroki Nakahara

    ICFPT 2007: INTERNATIONAL CONFERENCE ON FIELD-PROGRAMMABLE TECHNOLOGY, PROCEEDINGS 217-223 2007

    DOI: 10.1109/FPT.2007.4439252  

  76. A soft error tolerant LUT cascade emulator Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao

    PROCEEDINGS OF THE 15TH ASIAN TEST SYMPOSIUM 115-+ 2006

    ISSN: 1081-7735

  77. A fast logic simulator using a look up table cascade emulator Peer-reviewed

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    ASP-DAC 2006: 11TH ASIA AND SOUTH PACIFIC DESIGN AUTOMATION CONFERENCE, PROCEEDINGS 466-472 2006

    DOI: 10.1109/ASPDAC.2006.1594729  

    ISSN: 2153-6961

  78. Realization of Sequential Circuits by Look-Up Table Rings

    Tsutomu Sasao, Hiroki Nakahara, Munehiro Matsuura, Yukihiro Iguchi

    Midwest Symposium on Circuits and Systems 1 I517-I520 2004

    ISSN: 1548-3746

Show all ︎Show first 5

Misc. 47

  1. アンサンブル学習を用いたスパースCNNのFPGA実装に関して—Many Universal Convolution Cores for Ensemble Sparse Convolutional Neural Networks—VLSI設計技術

    倉持 亮佑, 佐田 悠生, 下田 将之, 佐藤 真平, 中原 啓貴

    電子情報通信学会技術研究報告 = IEICE technical report : 信学技報 119 (371) 67-72 2020/01

    Publisher: 東京 : 電子情報通信学会

    ISSN: 0913-5685

  2. 畳み込みニューラルネットワークを用いた単眼深度推定のFPGA実装について—An FPGA Implementation of Monocular Depth Estimation—VLSI設計技術

    佐田 悠生, 下田 将之, 佐藤 真平, 中原 啓貴

    電子情報通信学会技術研究報告 = IEICE technical report : 信学技報 119 (371) 73-78 2020/01

    Publisher: 東京 : 電子情報通信学会

    ISSN: 0913-5685

  3. 意味的領域分割のための組み込みシステム向け疎な全畳み込みニューラルネットワークのFPGA実装の検討—Filter-wise Pruning Approach to FPGA Implementation of Fully Convolutional Network for Semantic Segmentation—VLSI設計技術

    下田 将之, 佐田 悠生, 中原 啓貴

    電子情報通信学会技術研究報告 = IEICE technical report : 信学技報 118 (430) 25-30 2019/01

    Publisher: 東京 : 電子情報通信学会

    ISSN: 0913-5685

  4. 特徴マップを空間分割したCNNのFPGAにおける小メモリ実装—Spatial-Separable Convolution : Low memory CNN for FPGA—VLSI設計技術

    神宮司 明良, 下田 将之, 中原 啓貴

    電子情報通信学会技術研究報告 = IEICE technical report : 信学技報 118 (457) 7-12 2019

    Publisher: 東京 : 電子情報通信学会

    ISSN: 0913-5685

  5. 全2値化畳み込みニューラルネットワークとそのFPGA実装について : FPT2017デザインコンテスト参加報告—All Binarized Convolutional Neural Network and Its implementation on an FPGA : FPT2017 Design Competition Report

    下田 将之, 佐藤 真平, 中原 啓貴

    電子情報通信学会技術研究報告 = IEICE technical report : 信学技報 117 (378) 7-11 2018/01

    Publisher: 東京 : 電子情報通信学会

    ISSN: 0913-5685

  6. FPGA向けディープラーニング開発環境GUINNESSについて—GUINNESS : A GUI based Binarized Deep Neural Network Framework for an FPGA

    中原 啓貴, 米川 晴義, 藤井 智也, 下田 将之, 佐藤 真平

    電子情報通信学会技術研究報告 = IEICE technical report : 信学技報 117 (221) 51-56 2017/09

    Publisher: 東京 : 電子情報通信学会

    ISSN: 0913-5685

  7. 依頼講演 BRein Memory : バイナリ・インメモリ再構成型深層ニューラルネットワークアクセラレータ (集積回路)

    安藤 洸太, 植吉 晃大, 折茂 健太郎, 米川 晴義, 佐藤 真平, 中原 啓貴, 池辺 将之, 浅井 哲也, 高前田 伸也, 黒田 忠広, 本村 真人

    電子情報通信学会技術研究報告 = IEICE technical report : 信学技報 117 (167) 101-106 2017/07/31

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  8. 畳み込みニューラルネットワークの全2値化に関する一検討—Consideration of All Binarized Convolutional Neural Network

    下田 将之, 藤井 智也, 米川 晴義, 佐藤 真平, 中原 啓貴

    電子情報通信学会技術研究報告 = IEICE technical report : 信学技報 117 (153) 131-136 2017/07

    Publisher: 東京 : 電子情報通信学会

    ISSN: 0913-5685

  9. 依頼講演 BRein Memory : バイナリ・インメモリ再構成型深層ニューラルネットワークアクセラレータ (情報センシング)

    安藤 洸太, 植吉 晃大, 折茂 健太郎, 米川 晴義, 佐藤 真平, 中原 啓貴, 池辺 将之, 浅井 哲也, 高前田 伸也, 黒田 忠広, 本村 真人

    映像情報メディア学会技術報告 = ITE technical report 41 (25) 101-106 2017/07

    Publisher: 映像情報メディア学会

    ISSN: 1342-6893

  10. A Binarized Deep Neural Network for an Embedded System

    61 8p 2017/05/23

    Publisher: システム制御情報学会

  11. Accelerated Ternarized Deep Neural Network by sparse matrix calculation

    117 (46) 7-11 2017/05/22

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  12. A Memory Reduction with Neuron Pruning for a Convolutional Neural Network : Its FPGA Realization

    116 (417) 55-60 2017/01/23

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  13. Implementation of Binarized Deep Neural Network for FPGA Considering Power Performance Enhancement

    116 (417) 127-132 2017/01/23

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  14. A Memory Based Realization of the Binarized Deep Convolutional Neural Network

    116 (210) 63-68 2016/09/05

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  15. A Realization of Deep Convolutional Neural Network using the Nested RNS on an FPGA including the Constant Division

    115 (400) 227-232 2016/01/19

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  16. A Realization of Deep Convolutional Neural Network using the Nested RNS on an FPGA including the Constant Division

    115 (398) 227-232 2016/01/19

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  17. An FFT Circuit Using Nested RNS in a Digital Spectrometer for a Radio Telescope

    115 (343) 39-44 2015/12/01

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  18. A Design Method Using Discrete Particle Swarm Optimization for a Deep Neural Network Based on Nested RNS

    115 (228) 63-68 2015/09/18

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  19. A Deep Convolutional Neural Network Based on Nested Residue Number System

    115 (109) 91-96 2015/06/19

    Publisher: 電子情報通信学会

    ISSN: 0913-5685

  20. An AWF Digital Spectrometer for a Radio Telescope

    NAKAHARA Hiroki, NAKANISHI Hiroyuki, IWAI Kazumasa

    Technical report of IEICE. VLD 114 (426) 67-72 2015/01/29

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    A radio telescope analyzes radio frequency (RF) received from celestial objects. It consists of an antenna, a receiver, and a spectrometer. The spectrometer converts the time domain into the frequency domain by a FFT operation. In the spectrometer, first, it multiples the window coefficient to the received signal. Second, it applies the FFT operation. Third, it converts to the absolute of a complex number. Finally, to reduce the noise, it accumulates obtained power spectrum. We call this a WFA spectrometer. Since an AD converter is faster than the FPGA, a parallel FFT computation is desired. However, since the amount of hardware for the FFT becomes bottleneck, the conventional WFA does not realized the high-performance analysis. This paper proposes an AWF spectrometer which replaces the order of operations. Since the AWF spectrometer reduces the parallelism of the FFT, it is smaller than the conventional WFA spectrometer. Also, the paper proposes a off-chip memory realization which is highly efficient use of the FPGA. Experimental results show that the proposed AWF spectrometer outperforms conventional spectrometers.

  21. A High-Speed FFT for a Solar Radio Burst Observation On a Radio Telescope

    NAKAHARA Hiroki, CHISHIKI Yohei, IWAI Kazumasa, NAKANISHI Hiroyuki

    IEICE technical report 113 (325) 1-6 2013/11/27

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    A radio telescope analyzes radio frequency (RF) received from celestial objects. It consists of an antenna, a receiver, and a spectrometer. The spectrometer converts the time domain into the frequency domain by a FFT operation. A solar radio burst observation requires a high-speed FFT. This paper proposes the high-speed FFT based on Six-Step FFT algorithm. We implement P parallel FFT based on Six-Step FFT algorithm on the Xilinx Virtex 7 VC707 evaluation board. Experimental results shows that the proposed parallel FFT outperforms conventional FFTs.

  22. An Update Method for a CAM Emulator using a LUT Cascade Based on an EVBDD

    KUSHIYAMA Kensuke, NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    IEICE technical report 113 (325) 7-12 2013/11/27

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    The core routers forward packets by IP-lookup using longest prefix matching (LPM). With the rapid growth of the Internet, LPM has become the bottleneck in network traffic management. We have proposed an area-efficiency and high-performance LPM architecture using a LUT cascade based on an edge-valued binary decision diagram (EVBDD). As for the internet, the registered vector is frequency updated. This paper proposes an algorithm for the update of the LUT cascade. Its update time is O(n), where n is the length of the registered vector. We implemented the proposed algorithm on the ARM processor of the Zynq-FPGA. Experimental shows that, as for the normalized area and lookup speed, our architecture outperforms existing FPGA realizations.

  23. An Architecture for IPv6 Lookup Using Parallel Index Generation Units

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    IEICE technical report. Computer systems 112 (376) 25-30 2013/01/16

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    This paper shows an area-efficiency and high-performance architecture for the IPv6 lookup using parallel index generation units (IGUs) and a priority encoder. To reduce the size of memory for the IGU, we adopt a liner transform and a row-shift decomposition. Also, this paper shows a design method for parallel IGUs with given prefixes. Experimental shows that, as for the normalized area and lookup speed, our architecture outperforms existing FPGA realizations.

  24. AS-1-3 On a Muti-Valued Processor Based on a Decomposed MTMDDs for CF

    Nakahara Hiroki, Sasao Tsutomu, Matsuura Munehiro

    Proceedings of the Society Conference of IEICE 2012 "S-5"-"S-6" 2012/08/28

    Publisher: The Institute of Electronics, Information and Communication Engineers

  25. On a Wideband Fast Fourier Transform Using A Piecewise Linear Approximation : Applied to a Radio Telescope Spectrometer

    NAKAHARA Hiroki, NAKANISHI Hiro, SASAO Tsutomu

    Technical report of IEICE. VLD 112 (114) 43-48 2012/06/25

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    In a radio telescope, the spectrometer analyzes the radio frequency (RF) received from celestial objects at the frequency domain by performing the fast fourier transform (FFT). The FFT can be realized by the radix 2^k FFT (R2^k FFT). In Radio Astronomy, the number of points for the FFT is larger than that for the general purpose one. Thus, the twiddle factor memory is too large to implement. In this paper, we implement the twiddle factor by the piecewise linear approximation circuit consisting of a small memory, a multiplier, an adder, and a small logic circuit. We analyze the approximation error for the piecewise liner approximation circuits. We implemented the 2^<30> points FFT by the R2^k FFT with the piecewise linear approximation circuit. Compared with other FFT libraries, the R2^k FFT with the piecewise linear approximation using the memory is faster and smaller. Compared with the SETI spectrometer for 2^<27>-FFT in one second, the eight parallelized proposed ones for 2^<27>-FFT is 41.62 times faster, and that for 2^<30>-FFT is 5.20 times faster.

  26. On a Decomposed MTMDDs for CF Machine

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    Technical report of IEICE. VLD 111 (397) 31-36 2012/01/18

    Publisher: The Institute of Electronics, Information and Communication Engineers

    More details Close

    A decomposed multi-terminal multi-valued decision diagrams for characteristic function (MTMDDs for CF) represents decomposed circuits. A previous work shows that the decomposed MTMDDs for CF is smaller than monolithic decision diagrams for complex functions. This paper shows the decomposed MTMDDs for CF machine. First, we introduce the decomposed MTMDDs for CF. Then, we consider the instruction sets to evaluate the decomposed MTMDDs for CF. Next, we show the architecture for the decomposed MTMDDs for CF machine. We compare the decomposed MTMDDs for CF machine with other MPUs using MCNC benchmark functions. The decomposed MTMDDs for CF machine is 13.12 times faster than Altera's Nios II processor, and is 1.91 times faster than Intel's Atom N455 processor. As for the power-delay product, it is 60.84 times smaller than Nios II processor, and is 18.66 times smaller than Atom N455 processor.

  27. On a Power-Delay Product for a Heterogeneous MDD for ECFN Machine

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    IEICE technical report 111 (323) 1-6 2011/11/28

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    This paper analyzes a power-delay product for a HMDD for an ECFN (Heterogeneous Multi-valued Decision Diagram for Encoded Characteristic Function for Non-zero outputs) Machine that emulates the HMDD for ECFN. First, we introduce the HMDD for ECFN representing the multi-output logic function. Then, we show an architecture for the HMDD for ECFN machine. Next, we obtain the delay time and the power consumption for the HMDD for ECFN machine using the MCNC benchmark function. Finally, we analyze the power-delay product for the HMDD for ECFN machine. Compared with the Intel's Core i5, whose clock frequency is 2.4 GHz, as for the delay time, the HMDD for ECFN machine is 1.40-4.27 times shorter than Core i5, and as for the power-delay product, it is 15.1-46.6 times smaller.

  28. A Regular Expression Matching Circuit Based on a Decomposed Automaton

    2010 (5) 6p 2011/02

    Publisher: 情報処理学会

    ISSN: 1884-0930

  29. On a Prefetching Heterogeneous MDD Machine

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    IEICE technical report 110 (319) 13-18 2010/11/23

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    This paper shows a heterogeneous multi-valued decision diagram machine (HMDDM). First, we introduce a standard heterogeneous multi-valued decision diagram (HMDD). Then, we show a method to prefetch index for the HMDDM. Next, we introduce a code generation method for the HMDDM utilizing given memory size efficiently. We implemented the standard HMDDM and the prefetching HMDDM on an FPGA. The implementation results show that the prefetching HMDDM is 18.2% faster than the standard HMDDM. Also, we compared with Intel's Core2Duo (1.2 GHz) and a quartary decision diagram machine (QDDM). As for the execusion time, the prefetching HMDDM is 9.57-11.85 times faster than the QDDM, and 16.22-20.08 times faster than the Core2Duo.

  30. A Regular Expression Matching Circuit Based on an NFA with Mutli-Character Consuming

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    IEICE technical report 110 (204) 13-18 2010/09/09

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    This paper shows an implementation of a regular expression circuit based on an NFA (Non-deterministic finite automaton). Also, it shows that the NFA based one is superior to the DFA (Deterministic finite automaton) based one, with respect to the area complexity and the time complexity. A regular expression matching circuit is produced as follows: First, the given regular expressions are converted into a non-deterministic finite automaton (NFA). Then, to reduce the number of states, the NFA is converted to a modular non-deterministic finite automaton (MNFA(p)) with p-character-consuming transition. Finally, a finite-input memory machine (FIMM) to detect p-characters is generated, and the matching elements (MEs) realizing the states for the MNFA(p) are generated. We designed MNFA(p) for different p on Xilinx FPGA. As for the performance per area, our method is 6.2-18.6 times better than the DFA-based methods, and is 1.8 times better than the NFA-based method. Then, we derive an optimal value p that efficiently uses both LUTs and embedded memories of the FPGA.

  31. A Quaternary Decision Diagram Machine: Optimization of Its Code

    Tsutomu Sasao, Hiroki Nakahara, Munehiro Matsuura, Yoshifumi Kawamura, Jon T. Butler

    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS E93D (8) 2026-2035 2010/08

    DOI: 10.1587/transinf.E93.D.2026  

    ISSN: 1745-1361

  32. A Parallel Branching Program Machine for Sequential Circuits: Implementation and Evaluation

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura, Yoshifumi Kawamura

    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS E93D (8) 2048-2058 2010/08

    DOI: 10.1587/transinf.E93.D.2048  

    ISSN: 0916-8532

  33. A Comparison of Two Approximate String Matching Algorithms Implemented on an FPGA

    SHIMIZU Keisuke, NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    IEICE technical report 109 (462) 145-150 2010/03/03

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    An approximate string matching finds the most similar pattern in the text. A dynamic programming is used for the approximate string matching. This paper considers two algorithms: Naive method and LL method. Also, it derives area complexities with respect to the pattern length. Our implementations on an Altera's FPGA agree with the derived area complexities.

  34. A Virus Scanning Engine Using a Parallel Sieve Method and the MPU

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro, KAWAMURA Yoshifumi

    IEICE technical report 109 (320) 25-30 2009/11/26

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    In this paper, we show a new architecture for the virus scanning machine, which is different from that of the intrusion detection machine. The proposed method uses the two-stage matching, which is area-throughput efficient. That is, in the first stage, the hardware filter quickly scans to find possible matches, and in the second stage, the MPU scans the real match by a brute-force method. To make the hardware filter simply, we will introduce finite-input memory machine(FIMM). To reduce the memory size in the FIMM, we will introduce the parallel sieve method. The proposed method uses memory, so the power consumption is lower than the TCAM-based method. The system is implemented on the Stratix III FPGA and three off chip SRAMs, where all ClamAV virus patterns(514287)are stored. Comparison with existing methods, as for the area-throughput ratio, our method is 1.41-31.36 times more efficient.

  35. Emulation of Sequential Circuits by a Parallel Branching Program Machine

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro, KAWAMURA Yoshifumi

    IEICE technical report 108 (478) 111-116 2009/03/04

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    The parallel branching program machine (PBM128) consists of 128 branching program machines (BMs) and a programmable interconnection. To represent logic functions on BMs, we use quaternary decision diagrams. To evaluate functions, we use 3-address quaternary branch instructions. We realized many benchmark functions on PBM128, and compared its memory size and computation time with the Intel's Core2Duo microprocessor. PBM128 requires approximately quarter of the memory for the Core2Duo, and is 21.4-96.1 times faster than the Core2Duo.

  36. A parallel branching program machine for emulation of sequential circuits

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura, Yoshifumi Kawamura

    Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 5453 261-267 2009

    DOI: 10.1007/978-3-642-00641-8_26  

    ISSN: 0302-9743 1611-3349

  37. A Method of Design and Update for An Address Generator Using a Hybrid Method

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    2008 (2) 73-78 2008/01/16

    Publisher: Information Processing Society of Japan (IPSJ)

    ISSN: 0919-6072

    More details Close

    An address table relates k different registered vectors to the indices from 1 to κ. An address generation function represents the address table. This paper presents a realization of an address generation function with a hybrid method using a hash memory and a look-up table(LUT) cascade. The amount of hardware of the hybrid method is shown. Also, an update method for registered vectors is presented. We compared three different realizations: the hybrid method, CAMs produced by the Xilinx Core Generator, and the multiple LUT cascades. Experimental results show that the area for hybrid method is only 8 to 12% of the area for Xilinx CAMs, and is 35% of area for the multiple LUT cascades. Although our update method is complicated, the hybrid method requires smaller area and faster than conventional methods.

  38. A Method of Design and Update for An Address Generator Using a Hybrid Method

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    IEICE technical report 107 (418) 73-78 2008/01/16

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    An address table relates k different registered vectors to the indices from 1 to κ. An address generation function represents the address table. This paper presents a realization of an address generation function with a hybrid method using a hash memory and a look-up table(LUT) cascade. The amount of hardware of the hybrid method is shown. Also, an update method for registered vectors is presented. We compared three different realizations: the hybrid method, CAMs produced by the Xilinx Core Generator, and the multiple LUT cascades. Experimental results show that the area for hybrid method is only 8 to 12% of the area for Xilinx CAMs, and is 35% of area for the multiple LUT cascades. Although our update method is complicated, the hybrid method requires smaller area and faster than conventional methods.

  39. A Method to Evaluate Logic Functions Based On Decison Diagram Using Memory Packing

    TANAKA Hiroyuki, NAKAHARA Hiroki, MATSUURA Munehiro, SASAO Tsutomu

    IEICE technical report 106 (552) 45-50 2007/03/02

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    This paper proposes a method to evaluate logic functios using decision diagrams. Quasi-Reduced Multi-valued Decision Diagrams(QRMDDs) are used for logic simulation. This paper also shows a method to reduce memory requirement and computation time by memory packing. The speedup is due to the reduction of cache miss.

  40. A PC-based logic simulator using a look-up table cascade emulator

    Hiroki Nakahara, Tsutomu Sasao, Munehiro Matsuura

    IEICE TRANSACTIONS ON FUNDAMENTALS OF ELECTRONICS COMMUNICATIONS AND COMPUTER SCIENCES E89A (12) 3471-3481 2006/12

    DOI: 10.1093/ietfec/e89-.a.12.3471  

    ISSN: 1745-1337

  41. A Soft-Error Tolerant Look-Up Table Cascade Emulator

    NAKAHARA Hiroki, SASAO Tsutomu

    IEICE technical report 106 (198) 7-11 2006/07/25

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    An LUT cascade emulator realizes an arbitrary sequential circuit. We convert the combinational part into multiple LUT cascades, and store LUT(cell) data into a memory in the LUT cascade emulator. It evaluates multi-output logic functions by reading cell data sequentially. To improve torelance to soft errors, we encode cell data stored in the memory by error correcting codes. Also, we add an error-correcting circuit. A scanning circuit periodically scans the memory. When it detects a soft error, it remove the error by write-backing the correct data into the memory. To avoid the soft error for flip-flops, we employ a TMR (Triple Module Redundancy) technique. Our method detects the soft errors in a single bit. Also, the mission time of our method is more than 1000x of an ordinary LUT cascade emulator.

  42. A memory-based programmable logic device using look-up table cascade with synchronous static random access memories

    Kazuyuki Nakamura, Tsutomu Sasao, Munehiro Matsuura, Katsumasa Tanaka, Kenichi Yoshizumi, Hiroki Nakahara, Yukihiro Iguchi

    JAPANESE JOURNAL OF APPLIED PHYSICS PART 1-REGULAR PAPERS BRIEF COMMUNICATIONS & REVIEW PAPERS 45 (4B) 3295-3300 2006/04

    DOI: 10.1143/JJAP.45.3295  

    ISSN: 0021-4922

  43. A Fault Tolerant Look-Up Table Cascade Emulator

    NAKAHARA Hiroki, SASAO Tsutomu

    IEICE technical report 105 (647) 31-36 2006/03/03

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    An LUT cascade emulator realizes an arbitrary sequential circuit. We convert the combinational part into multiple LUT cascades, and store LUT (cell) data into a memory in the LUT cascade emulator. It evaluates multi-output logic functions by reading cell data sequentially. To improve availability of the LUT cascade emulator, we add a self-checking circuit. Cell data stored in the memory are encoded into Berger codes. The self-checking circuit watches for whether memory outputs are Berger codes. The self-checking circuit can detect uni-directional errors of the memory for logic, a single stuck-at fault of a decoder, and a single stuck-at fault of the programmable connection circuit for rails. When a fault is detected, the monitor stops the LUT cascade emulator, and goes to the fault avoidance mode. First, it diagnoses the fault of the LUT cascade emulator. Next, it rewrites memory contents by memory-packing that avoids the failure parts of the memory for logic. Therefore, our method can improve the availability of the circuit.

  44. A design algorithm for sequential circuits using LUT rings

    H Nakahara, T Sasao, M Matsuura

    IEICE TRANSACTIONS ON FUNDAMENTALS OF ELECTRONICS COMMUNICATIONS AND COMPUTER SCIENCES E88A (12) 3342-3350 2005/12

    DOI: 10.1093/ietfec/e88-a.12.3342  

    ISSN: 1745-1337

  45. A Logic Simulation using an Look-Up Table Cascade Emulator

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    2005 (121) 185-190 2005/11/30

    Publisher: Information Processing Society of Japan (IPSJ)

    ISSN: 0919-6072

    More details Close

    This paper shows a cycle-based logic simulation method using an LUT cascade emulator. The LUT cascade emulator is an architecture that emulates LUT cascades, where multiple-output LUTs (cells) are connected in series. The LUT cascade emulator has a control part, a large memory, and registers. It connects the memory to each register by a programmable interconnection circuit, and evaluates the given circuit stored in the memory. This method realizes the LUT cascade emulator on a PC by a software. Experimental results show that this method is 18-2621 times faster than the LCC.

  46. A Memory-Based Programmable Logic Device Using a Look-Up Table Cascade with Synchronous SRAMs

    NAKAMURA Kazuyuki, SASAO Tsutomu, MATSUURA Munehiro, TANAKA Katsumasa, YOSHIZUMI Kenichi, NAKAHARA Hiroki, IGUCHI Yukihiro

    2005 314-315 2005/09/13

  47. A Design Algorithm for Sequential Circuit Synthesis using LUT Ring

    NAKAHARA Hiroki, SASAO Tsutomu, MATSUURA Munehiro

    IEICE technical report. Dependable computing 103 (482) 145-150 2004/12/02

    Publisher: The Institute of Electronics, Information and Communication Engineers

    ISSN: 0913-5685

    More details Close

    This paper shows a design method for a sequential circuit by using a Look-Up Table(LUT) ring. The method consists of two steps: The first step partitions the outputs into groups. The second step realizes them by LUT cascades, and stores the content of cells in the cascades into the memory. We also compare the method with other methods to realize sequential circuits. With the presented algorithm, we can easily design sequential circuits satisfying given specifications.

Show all ︎Show first 5

Research Projects 7

  1. Development of a Binary Vision Transformer Hardware

    Offer Organization: Japan Society for the Promotion of Science

    System: Grants-in-Aid for Scientific Research

    Category: Grant-in-Aid for Scientific Research (B)

    Institution: Tohoku University

    2024/04/01 - 2029/03/31

  2. On a noise convolutional neural network

    Offer Organization: Japan Society for the Promotion of Science

    System: Grants-in-Aid for Scientific Research

    Category: Grant-in-Aid for Scientific Research (B)

    Institution: Tokyo Institute of Technology

    2019/04/01 - 2024/03/31

  3. FPGA向けAIアクセラレータの開発

    中原 啓貴

    Offer Organization: 科学技術振興機構

    Category: 産学が連携した研究開発成果の展開/研究成果展開事業/研究成果最適展開支援プログラム(A-STEP)/実装支援/返済型

    2024 -

    More details Close

    画像処理において、ニューラルネットワーク処理で特徴を抽出し定量化することにより、画像に写る特定の人・物を認識・検出できる。 しかし、高精度に検出するためには、特徴の抽出に多くのパラメーターが必要となり、AIを作動させるためのメモリー量や消費電力が肥大し、コストが増大する。一方で、パラメーター数を減らすと検出精度が劣化するというジレンマがあり、適用できるアプリケーションが限定的であった。 Tokyo Artisan Intelligence株式会社は、バッチ正規化回路を不要とするニューラルネットワーク回路により、消費電力およびコストを抑えつつ、高精度検出を可能にする技術を活用し、映像処理AIシステムを開発してきた。 本開発では、FPGA向けAI処理専用カスタムアクセラレータを開発する。消費電力およびコストを低減しながら、高い検出精度を保つAIを利用可能となり、また、顧客ごとのカスタマイズ性が高まり、将来的には、大量のデータ処理が求められるインフラ点検や工場における人体接近検知の自動化への展開も期待される。

  4. Development of Next Generation Spectrometer for Radio Telescope

    NAKAHARA HIROKI

    Offer Organization: Japan Society for the Promotion of Science

    System: Grants-in-Aid for Scientific Research

    Category: Grant-in-Aid for Young Scientists (A)

    2015/04/01 - 2019/03/31

    More details Close

    We implemented an algorithm in which the operation order of spectrometers has been changed and an FFT circuit based on Residue Number System (RNS) is applied to the existing FPGA board (ROACH2 board), which is an existing facility. We compared it with the existing spectrometer released by CASPER (The Collaboration for Astronomy Signal Processing and Electronics Research). A 50 times wider and 2^16 points resolution spectrometer was realized by our development technologies. The data classifier after observation was realized for a CNN (Convolutional Neural Network). We reduced the size of CNN hardware by binary precision and sparse (Ternary, pruning zero weights) and clarified the practicability of FPGA implementation.

  5. 超省電力、低コスト組込みシステム人工知能モジュールの開発

    中原 啓貴

    Offer Organization: 科学技術振興機構

    Category: 産学が連携した研究開発成果の展開/研究成果展開事業/研究成果最適展開支援プログラム(A-STEP)/シーズ育成タイプ

    2017 -

  6. A general purpose processor based on a multi-valued decision diagram

    NAKAHARA Hiroki

    Offer Organization: Japan Society for the Promotion of Science

    System: Grants-in-Aid for Scientific Research

    Category: Grant-in-Aid for Young Scientists (B)

    2012/04/01 - 2015/03/31

    More details Close

    We proposed the Multi-terminal multiple-valued decision diagram for characteristic function representing cluster decomposition (MTMDD for CF) as a new kind of a decision diagram, then presented at the international conferences. Also, we applied the edge-valued MDD(k), which is a variation of the MTMDD for CF, to realize a new embedded processor. We used it to the multi-core processor to realize the packet classification. Also, we developed the commercial packet classifier with co-development company.

  7. FPGAを用いた応用 組込みシステム(主にネットワーク機器) Competitive

    System: JST地域イノベーション創出総合支援事業

    2008 - 2012

    More details Close

    FPGAを用いた機器の研究開発。特に、ネットワークのセキュリティ・制御機器に関する高性能・低消費電力システムについて。

Show all Show first 5