アプリとサービスのすすめ

アプリやIT系のサービスを中心に書いていきます。たまに副業やビジネス関係の情報なども気ままにつづります

4 th AI Edge Contest Report


I took part in SIGNATE's 4th AI Edge Contest, so I wanted to write up a report/log of the experience.

It wasn't just a machine learning competition — the hardware side was just as serious.

Table of Contents
1. About the network
2. C++ application code optimizations
3. About the hardware platform )

1. About the network

1.1 Model and strategy used



I used U-Net, since it's easy to customize. For libraries, I used Keras and TensorFlow, with the following versions for the conversion work prior to quantization:

・Keras==2.2.4
・tensorflow-gpu==1.13.1

I chose U-Net because it's easy to customize the network — both for going from pretrained to fine-tuned, and for tweaking the architecture to improve accuracy.

Depth was set to 512. I kept the model size relatively small so processing speed wouldn't suffer, resulting in a memory footprint of **14,067,237**.

My strategy with this model was to beat the benchmark.

The reasoning: if I could clear the benchmark this way, it would create a meaningful differentiation from other participants in terms of ingenuity and processing speed, giving me an advantage.

Model size comparison with YOLOv3

Yolov3 62,002,753
Yolov3-tyny 8,861,918
This U-Net 14,067,237

Size comparison between depth 512 and depth 1024

Depth  Size
512 14,067,237
1024 31,055,557



Network diagram of the U-Net used this time


I also considered another option: using depth 1024 (memory size: 31,055,557) combined with lightweighting techniques such as pruning.

Also, since I expected this contest's segmentation task and preprocessing to increase the computational load on the hardware's PS side, I used softmax as the final layer.

Thanks to this, I was able to offload work to the softmax computation IP on the hardware side, increasing DPU utilization.

Approach I did not take


The approach I decided against was building a model with depth 1024 or greater and lightweighting it via pruning or distillation.

Techniques like pruning and distillation are heavily hardware-specific, so the learning curve and time/development cost would have been too high (especially self-taught, within the time available).

Without lightweighting, a depth-1024 model would exceed 30,000,000 parameters in memory, which directly and noticeably impacts processing speed — so I didn't go with this strategy.


1.2 Network design considerations to avoid compile-time errors

Certain layer configurations caused accuracy degradation or errors right before/after quantization, so I excluded those and built the network accordingly.

Improvements made

1. Strictly following the "Conv2D → BatchNormalization (BN) → ReLU" layer order


An invalid configuration is "ReLU → BN," which causes a compile-time error.

x = Conv2D(filters = n_filters, kernel_size = (kernel_size, kernel_size),
            kernel_initializer = 'he_normal', padding = 'same')(x)
x = BatchNormalization()(x)
x = Activation('relu')(x)

2. Using Dropout on the decoder side


Using BN on the decoder side (where Concatenate or Add layers are used) leads to accuracy degradation after quantization.

3. Using Conv2DTranspose instead of Conv2D right before softmax


Since I used softmax, in order to quantize the model I needed to use a layer other than Conv2D right before softmax (e.g., Conv2DTranspose, SeparableConv2D, etc.).

1.3 Using softmax with the DPU in mind

This time, to make use of the softmax computation IP on the hardware, I used softmax as the final layer of the U-Net as well.

Using softmax provided the following benefits:

• It increases DPU utilization

• It allows PL, DPU computation, and softmax computation to run in parallel across three threads

• Softmax gives higher accuracy than sigmoid or ReLU

Conditions for using softmax with U-Net


Through trial and error using the DPU in Vitis, I found the following:

• As a compile-time constraint, the layer immediately before softmax must be something other than Conv2D (e.g., Conv2DTranspose, SeparableConv2D, etc.)

# Code near the final layer of the U-Net (model) during fine-tuning
x=model.get_layer(index=-5).output
x = Conv2DTranspose(nClasses, kernel_size=1, use_bias=False)(x)
x = (Activation("softmax"))(x)


• If you don't specify `use_bias=False` for Conv2DTranspose, the output of the DNNDK library's `dpuGetOutputTensorScale()` changes, which can cause errors in the softmax output


1.4 Techniques used to improve accuracy

At depth 512, simply building the network wasn't enough to exceed IoU=60% — I needed accuracy-improvement techniques beyond just the network architecture itself.

Here are the main techniques I used.


Resizing that preserves the original image's (height, width) aspect ratio



=> I resized to shape (400, 680), preserving the original image's size ratio, and used OpenCV's resize while maintaining the aspect ratio. This improved prediction accuracy for small objects (signals, pedestrians) compared to resizing to a square shape like (224, 224).

I think this is because it reduced the loss of positional information during resizing.




Preprocessing to brighten dark images (mean pixel value below 80) via histogram equalization


=> This slightly improved accuracy for fine details in dark images. Since dark images tend to have skewed pixel distributions, I treated
"low mean pixel value = dark image"

and applied histogram equalization to images with a mean pixel value under 80 to brighten them.

def clahe(bgr):
    #plt.imshow(bgr),plt.show()
    lab = cv2.cvtColor(bgr, cv2.COLOR_BGR2LAB)
    #plt.imshow(lab),plt.show()
    lab_planes = cv2.split(lab)
    clahe = cv2.createCLAHE(clipLimit=6.0,tileGridSize=(8,8))
    lab_planes[0] = clahe.apply(lab_planes[0])
    lab = cv2.merge(lab_planes)
    return cv2.cvtColor(lab, cv2.COLOR_LAB2BGR)

def NormalizeImageArr(path, H, W):
    NORM_FACTOR = 255
    img = cv2.imread(path, 1)
    img = cv2.resize(img, (H, W), interpolation=cv2.INTER_NEAREST)
    if img.mean()<80:
        img = clahe(img)
    img = img.astype(np.float32)
    img = img/NORM_FACTOR
    return img

Training with augmentation


Augmentations that didn't change object position (noise, contrast adjustments, horizontal flip) were the most effective.
Vertical flip backfired when objects (like cars) never appear upside down.

Also, rather than applying all augmentations at once, gradually introducing them step by step seemed to steadily improve accuracy.
I trained with augmentation using the following steps:

Epoch Dataset IOU Augmentation
100 CitySpacuies None  None
200 train: 2,143 images (contest images), val: 100 images (contest images) train=83.8%、Val = 74% None
200 train: 2,143 images (flipped only), val: 100 images (flipped only) train=76%、Val = 68% Horizontal flip (also applied to validation)
100 train:4286 images、val:100 images train=88.5%、Val = 75.8% Horizontal flip (not applied to val)
100 train:4286 images、val:100 images train=89.2%、Val = 77.5% contrast-based(not applied to val)


Pretraining on the Cityscapes dataset


I borrowed this idea from a precedent set in a previous contest, and it improved accuracy considerably.

I also tried adding residual structures and sub-layers favorable for segmentation, but at depth 512 the model's representational capacity had reached its limit, so this had almost no effect. It was also tough that layers like SENet (which use multiply operations) ran into compile-time errors, which was a hard constraint that prevented further accuracy improvements in that direction.

What I learned through this PDCA process is that applying techniques haphazardly, without some hypothesis or logic for why they should improve accuracy, mostly ends up being wasted effort.




1.5 Final network results (IoU, etc.)

In the end, with a model size of 14,067,237, I achieved an IoU of roughly 61%.

2. C++ application code optimizations

To maximize processing speed and squeeze out hardware performance, I focused especially on two things in C++.

2.1 Reducing computational load

Since the PS-side computation was heavy this time, and I used three threads, cutting down and simplifying redundant code had a meaningful effect on processing speed.

In particular, the following kinds of rewrites sped things up by roughly 30ms per improvement:

・Making `for` loops more efficient

・Removing unnecessary functions and unnecessary calls to those functions

・Turning fixed values into constants





2.2. Three-way multithreading to draw out hardware performance

Since this contest involved a lot of PS-side computation outside of the DPU (e.g., preprocessing and `for` loops for segmentation), splitting the work across three threads sped things up by roughly 30–50ms.

Below is the difference in speed between two threads and three threads, under DPU parameters (B1152, "DSP48 USAGE=LOW", etc.):

Number of threads Average processing time per image (ms)
2 1061
3 1007




2.3 Further speed improvement by running PS, PL (DPU computation), and softmax computation in parallel across three threads

The original purpose of multithreading is to speed things up by running PS and PL in parallel. However, since I was using the DNNDK library this time, PL work is split into two independent methods:

• DPU computation
• Softmax computation

DPU computation method dpuRunTask()
Softmax computation method dpuRunSoftmax()

Because of this, I made the following three things the targets of parallel multithreaded processing:

・ PS computation
・ DPU computation
・ Softmax computatio

By applying asynchronous processing std::lock_guard lock(mtx_) to both the DPU computation and softmax computation, I was able to run DPU computation and softmax computation in parallel, achieving a further speed improvement through multithreading.


Excerpt of the multithreaded function (main_thread()) that runs PS, DPU computation (PL), and softmax computation (PL) in parallel (key sections only)

#include <thread> 
#include <opencv2/opencv.hpp>
#include <opencv2/core.hpp>
#include <dnndk/dnndk.h>
#include <mutex>  
std::mutex mtx_;
〜〜
〜〜

int main_thread(DPUKernel *kernelConv, int s_num, int e_num, int tid){
  assert(kernelConv);
  DPUTask *task = dpuCreateTask(kernelConv, DPU_MODE_NORMAL); 
  〜〜〜
  // Main Loop
  int cnt=0;
  for(cnt=s_num; cnt<=e_num; cnt+=BLOCK_SIZE){
      for(int i=0; i<BLOCK_SIZE;i++){
        if(cnt+i>e_num) break;
        Mat img;
        resize(input_image[i], img, for_resize, INTER_NEAREST);
        // pre-process with histgram avaraving
        Mat clahe_img = img;
        if((int)mean(img)[0] < 80) {
           clahe_img = clahe_preprocess(img);	
        }
    
        float *softmax = new float[outWidth*outHeight*outChannel]
        // Set image into Conv Task with mean value
        set_input_image(task, outWidth, clahe_img);
        {
          std::lock_guard<std::mutex> lock(mtx_);
          dpuRunTask(task);
        }
        {
          std::lock_guard<std::mutex> lock(mtx_);
          //cout << "outScale : " << outScale << endl;
          int8_t *outAddr = (int8_t *)dpuGetOutputTensorAddress(task, CONV_OUTPUT_NODE);
          dpuRunSoftmax(outAddr, softmax, outChannel,outSize/outChannel, outScale);
        }

        // Post process
        PostProc(softmax, outHeight, outWidth, outChannel, image_file_name[i].c_str());
        delete[] softmax;
      }
  }
  dpuDestroyTask(task);
   return 0;
}


Three threads running PS, DPU computation, and softmax computation in parallel

3. About the hardware platform

3.1 Development environment

I referred to a Vitis-AI development environment article on Qiita. I didn't use the Vitis-AI-Runtime library, and instead built on top of the DNNDK library.


3.2 Notes on building the DPU hardware platform

I mainly referred to the Vitis-AI environment setup tutorial and materials from the 2nd AI Edge Contest (referred to below as "reference materials"), and built the platform by improving on the tutorial's platform.


3.2.1 Making use of the softmax computation IP


To make the most of DPU computation, I made use of the softmax computation IP.

Since I used Conv2DTranspose in the U-Net, and designed the model with this integration in mind, I used softmax computation.




Platform including softmax computation (unnecessary IPs already removed)

3.2.2 Building and refining the platform


Since development was based on DNNDK, I mainly refined the Vitis-AI platform tutorial, using the reference materials as a guide.

Initially, my first goal was to get a platform with a DPU installed working under the following conditions from the reference materials:

• A B1600 DPU integrated with softmax

•DPU frequency of 250MHz

To do this, I first took the tutorial's platform and:

1. Removed unnecessary IPs

2. Removed unnecessary clocks

and built the platform with WNS = 0.027 .

This time, because I needed to enable "DepthwiseConv" in order to use my model, I had to make changes to the parameters and frequency.


3.3 DPU parameters and frequency


That change made the parameter set too large — pushing the frequency higher caused the board to reboot, so I couldn’t get a DPU running with B1600 at 250MHz

On top of that, since “DepthwiseConv” uses 3,292 LUTs on B1600, I had to significantly reduce the DPU parameter resources compared to the reference materials.

In particular, since this model makes heavy use of convolutional layers, without enabling "Channel Augmentation," processing speed dropped considerably even with "DSP48 Usage" set to High. So I made the following mandatory:

・Enable "Channel Augmentation"

・Enable "DepthWiseConv"

As a result, at frequencies of 225MHz or higher with "DSP48 USAGE" set to HIGH, the board would reboot. So in the end, I settled on a frequency of 200MHz with the following parameters for the DPU:

Frequency 200MHz
DPU B1600(ReLU+ReLU6)
Channel Augmentation Enable
DepthWiseConv Enable
PoolAverage Disnable
DSP48 USAGE HIGH
RAM USAGE LOW
Softmax Enable

Given the resource constraints, I wasn't able to push the frequency any higher than this.


3.4 Improving processing speed via impl and synth strategies

At 200MHz with this DPU parameter set, the frequency alone wasn't enough to draw out the DPU's full performance. So I tried several strategy combinations to see if I could improve processing speed further, referring to an article from Fixstars' site.

According to the "High-Density FPGA Design Guide", the more resources you use, the lower the integration density needs to be in order to make effective use of those resources.

Since the DPU had a lot of resources allocated this time, I chose a combination of strategies that spread the logic out to lower the integration density — this sped things up by about 35ms:

impl Congestion_SpreadLogic_high
Synth Flow_AreaOptimized_high
WNS 0.131 ns

I also tried the SSI-distributing strategy "impl: Congestion_SSI_SpreadLogic_low." While SSI has lower power consumption, its integration density is higher, so the combination above gave better processing performance.

With this strategy combination, B1600, and a 200MHz frequency, I was able to fit a DPU that satisfied all the constraints.

Compared to PS, the utilization of DPU and softmax ended up as follows:

PS & PL Tototal 93%
DPU 43%
Softmax 6%

Power consumption report for this build

References



・vitis-AI platform site(qiita)

・DPU-TRD

・Xilinx GitHub Vitis-AI-TUTORIAL

・Trying different Vivado synthesis/implementation strategies(WNS & runtime)

・2nd AI Edge Contest materials

・Zynq DPU v3.2 Guide

・High-Density FPGA Design Guide

english translation problem(英作文)

NO.1



楽しいはずの海外旅行にもトラブルはつきものだ。たとえば,悪天候や自然災害 によって飛行機が欠航し,海外での滞在を延ばさなければならないことはさほど珍し いことではない。いかなる場合でも重要なのは,冷静に状況を判断し,当該地域につ いての知識や情報,さらに外国語運用能力を駆使しながら,目の前の問題を解決しよ うとする態度である。



example answer

If you travel abroad, you will have fun but also encounter accidents you don't expect. It isn't unusual that you have to stay in the foreign country longer than you have planned because of your flight schedule being canceled by sudden events such as natural disasters or bad weather. In such a case, it is important that you try to gather necessary information and knowledge about the place you are staying by using languages you can speak.


better answer

Trouble is part of traveling abroad, even though such trips are supposed to be enjoyable. For example, it is not particularly unusual to have to extend your stay abroad because your flight has been canceled due to bad weather or a natural disaster. In any situation, what matters is maintaining a calm attitude, assessing the situation carefully, and trying to solve the problem at hand by making full use of your knowledge and information about the local area, as well as your foreign-language skills.




NO2

人と話していて,音楽でも映画でも何でもいいが,何かが好きだと打ち明けると, たいていはすぐさま,ではいちばんのお気に入りは何か,ときかれることになる。こ の問いは,真剣に答えようとすれば,かなり悩ましいものになりうる。いやしくも映 画なり音楽なりの愛好家である以上,お気に入りの候補など相当数あるはずであり, その中から一つをとるには,残りのすべてを捨てねばならない。

example answer

When you tell someone that you like something, whether it is music, movies, or anything else, you are usually asked right away what your favorite is. If you try to answer that question seriously, it can be surprisingly difficult. If you are truly a fan of movies or music, you probably have quite a few possible favorites, and choosing just one means giving up all the others.


better answer

When you are talking with someone and mention that you like something, whether it is music, movies, or anything else, you will usually be asked right away what your favorite is. If you try to answer this question seriously, it can be quite difficult. If you are truly a fan of movies or music, you are bound to have quite a few candidates for your favorite, and choosing just one means giving up all the others


NO3

人間の性格は見かけよりも複雑なので,相手のことが完全に分かることなどある はずがない。とは言うものの,初対面の人物とほんの少し言葉を交わしただけで,そ の人とまるで何十年も前からつきあいがあったかのような錯覚に陥ることがある。こ うしたある種の誤解が,時として長い友情のきっかけになったりもする。



my answer

Because one’s character is more complicated than you might expect, it may be impossible to understand them. Even if you have only had a short chat with someone you’ve met for the first time, you may feel as if you have known them for a long time. This kind of misunderstanding sometimes becomes an opportunity for long relationships.



ChatGPT with smart expression

“Because one’s character is more complex than one might expect, it can be impossible to understand them. Even after only a short chat with someone you’ve just met, you may feel as though you’ve known them for a long time. This kind of misunderstanding can sometimes lead to long-lasting relationships.”



NO4

私の意見では, 現代の若者は性別を問わず自分で調理できることが大切である。料理をおいしく仕上げるためには豊かな想像力や手先の器用さが要求されるので, 心身の健康にとても良い。 食材に意識的になれば自然への関心も高まる。さらに, 料理で友人をもてなすことができると, あるいは人と共同して料理ができると, 絆が深まることは間違いない。

my answer

I think that young people today should learn to cook by themselves regardless gender. Cooking is good for your health mentally and pysically because in order to make delicious food you have to use complicated tool and think what will be needed for it. If you have oppotunities to think about cooking materials, you will become more interested in the nature. for you friends, If you can enjoy for your meals and cook with them, you can have make more strong relationship with them.



ChatGPT4 modify

I think that young people today should learn to cook by themselves regardless of gender. Cooking is good for your health, both mentally and physically, because in order to make delicious food you have to use complex tools and think about what ingredients you will need. If you have opportunities to think about cooking materials, you will become more interested in nature. Cooking can also strengthen your relationships with your friends. If you can enjoy your meals and cook with them, you can have more fun and bond with them.





ChatGPT4 with smart expression

“Cooking is a valuable skill that young people of any gender should learn. It benefits your health in multiple ways, as it requires you to use sophisticated tools and plan ahead for the ingredients you need. Cooking also sparks your curiosity about nature and the sources of your food. Moreover, cooking can enhance your social life, as you can share your meals and recipes with your friends and have a great time together.”



NO.5



今日,睡眠不足は見過ごせない問題となっている。 原因の一つは, 社会全体が深夜も多くの人が起きていることを想定して動いていることである。 照明器具の発達も我々の体内時計を狂わせているのかもしれない。 その一方で多くの学校や会社の始まる時間は変わっていない。 こうして睡眠不足が生まれやすくなり, 日中の集中力の低下を引き起こすのだ。



Today, lack of enough sleep is one of the serious problem, partly because in modern socity most people tend to stay up late at night for many reason. And their life cycle is so different from what it used to be because of technology advanece which let them be able to stay up late at night. On the other hands, most people must get up early for school or work and these system haven't change for a long time. That's why in modern socity most people tend to stay up late at night, resulting in decline of their daily concentration.



conversational way by ChatGPT4

“Many people don’t get enough sleep these days, and that’s a big problem. One reason is that they have to wake up early for school or work, but they also like to stay up late at night. There are so many things to do with the new technology we have. But this changes our natural rhythm of sleeping and waking up. It makes us less focused and alert during the day.”



comment from ChatGPT4

I have simplified some of the words and sentences in your paragraph to make it sound more like spoken English. Here are some tips for converting written English to spoken English:

Use shorter sentences and simpler words. Avoid long and complex sentences that might confuse the listener.
Use contractions like “don’t”, “can’t”, “we’ve”, etc. They make your speech sound more natural and informal.
Use words and phrases that show your attitude or emotion. For example, you can use “and that’s a big problem” or “but they also like to stay up late at night” to express your opinion or contrast two ideas.
Use words and phrases that connect your sentences and ideas. For example, you can use “one reason is that” or “but this changes our natural rhythm” to show the relationship between your sentences.

spanish problems for practice

Nivel 1(基礎〜中級:主語補完が課題)

1. 「あの店、いつの間にか潰れていた。」

直訳寄り:
Sin darme cuenta, aquella tienda ya había cerrado.

意訳寄り:
No sé cuándo fue, pero esa tienda terminó echando el cierre.

(echar el cierre=店を畳む、口語的で自然な表現)

2. 「歳を取るごとに、一年が早く過ぎるように感じる。」

直訳寄り:
A medida que envejezco, siento que el tiempo de un año pasa más rápido que cuando era joven.

意訳寄り:
Cuanto más viejo me hago, más rápido me parece que pasan los años.

Nivel 2(中級〜上級:比喩・抽象名詞)

3. 「言葉の壁というのは、実は語彙力の問題ではなく、沈黙を恐れる気持ちの問題だと思う。」

直訳寄り:
Creo que la barrera del idioma no es, en realidad, un problema de vocabulario, sino un problema de miedo al silencio.

意訳寄り:
Al final, la barrera del idioma no tiene tanto que ver con cuántas palabras sepas, sino con el miedo a quedarte callado.

4. 「旅先で出会う人との関係は、深いようでいて、実は驚くほど軽やかなものだ。」

直訳寄り:
Las relaciones que se forman con la gente que conoces viajando parecen profundas, pero en realidad son sorprendentemente ligeras.

意訳寄り:
Con la gente que conoces de viaje, sientes una conexión que parece profunda, aunque en el fondo es liviana y pasajera.

Nivel 3(上級:京大レベル、婉曲・二重否定)

5. 「本当に大事なことほど、人はなかなか言葉にしない。むしろ、どうでもいいことばかり饒舌に語る。」

直訳寄り:
Cuanto más importante es algo, menos se habla de ello. Al contrario, la gente habla mucho de lo que no le importa en absoluto.

意訳寄り:
Parece que cuanto más nos importa algo, más callado nos quedamos al respecto; en cambio, no paramos de hablar de tonterías.

6. 「異国で暮らすということは、自分の常識が世界の常識ではないと、繰り返し思い知らされることだ。」

直訳寄り:
Vivir en otro país te hace darte cuenta, una y otra vez, de que lo que es común para ti no lo es en absoluto en otros países.

意訳寄り:
Vivir en el extranjero te obliga, una y otra vez, a entender que lo que para ti es normal, en otros lugares puede no serlo en absoluto.

4-1
私は、どんな本でも、読む以上は、はじめからしまいまで、途中をとばさずに、その全部を読むことを理想としている。なかなか実際にはできないけれども、そうありたいと思っている。

直訳寄り:
Siempre que leo un libro, considero que lo ideal es leerlo todo, sin saltarme ninguna parte, de principio a fin. No siempre puedo hacerlo en la práctica, pero me gustaría poder lograrlo.

意訳寄り:
Cuando empiezo un libro, mi ideal es terminarlo entero, sin saltarme ni una sola parte, de la primera página a la última. No siempre lo consigo, pero es justamente eso lo que me gustaría poder hacer siempre.


4-2(原題: 1981年度〔3〕(2))

一つのことを必ずやりとげようと思うなら、そのほかのことがだめになるのを嘆いてはならないし、他人の嘲笑をも恥ずかしいと思ってはならない。多くの事を犠牲にしなければ、一つの大きな仕事が完成するはずがない。

直訳寄り:
Si de verdad quieres llevar a cabo algo, no debes lamentarte de que las demás cosas te salgan mal, ni debes avergonzarte de las burlas de los demás. Sin sacrificar muchas cosas, es imposible completar una gran obra.

意訳寄り:
Cuando uno está decidido a lograr algo, no tiene sentido lamentarse porque lo demás se le vaya de las manos, ni avergonzarse de que otros se rían de uno. Ninguna gran obra se completa sin sacrificar mucho por el camino.


4-3(原題: 「鯛の味」の文章)

鯛を食べたことのない人に、鯛の味を説明しろと言われたら、皆さんはどんな言葉を選びますか。おそらく、どんな言葉を用いても言い表わす方法がないでしょう。このように、たった一つの物の味でさえ伝えることができないのですから、言語というものは案外不自由なものであります。

想定される難所:

• 「〜しろと言われたら」= si te pidieran que + 接続法過去(仮定法の定番パターン)
• 「食べたことのない人」=経験の欠如を表す関係詞節(que nunca ha probado / que jamás ha comido)
• 「どんな言葉を用いても」= por más que uses / uses las palabras que uses(譲歩の慣用構文、京大英作文でも頻出)
• 「〜でさえ〜できない」= ni siquiera + 動詞
• 「案外〜である」= resulta ser sorprendentemente… / en realidad es más … de lo que parece

install ubuntu 22.04 to window PC via USB

1 erase Partition on sd-card

# erase FAT32 area in USB CARD
$ diskutil eraseDisk JHFS+ MYDISKNAME disk2

# check state
$ diskutil list

referring site
・MacOS diskutilコマンドを使ってGUIのディスクユーティリティでは見えないボリュームを削除する #Mac - Qiita

・How to Delete Partition on Mac? How to Remove It on a Hard Drive?

2 install ubuntu img to USB

download Server install image 「ubuntu-22.04.5-desktop-amd64」
1. unmount

$ diskutil unmountDisk /dev/disk2

2. use flash and install
in the case of CUI

# it takes some time. you can check how it proceeds by 「Ctrl + T」
$ sudo dd if=ubuntu-22.04.5-desktop-amd64.iso of=/dev/disk2 bs=1m

3 make ubuntu PC
How to put a Linux ISO onto a USB stick and make it bootable on a Mac — The Ultimate Linux Newbie Guide
enter 'BIOS' in window and install ubuntu
=> launch PC and push Enter 4 times or more and F1.
=> enter BIOS. security => secure boot => secure Boot 「ON」 Allow Microsoft 3rd party UEFI CA 「ON」
=> save and exit => push Enter 4 times or more and F12.
=> BOT menu => USB HDD~ (Verbatim STORE N GO)
youtu.be

4 after install
if you can succeed peacefully, erase the ubuntu img from USB.

$ diskutil eraseDisk JHFS+ MYDISKNAME disk2

if you leave it without doing anything, USB may be broken.

ros2のツール一覧まとめ[2025/01/16]

1. rosbag

# 必要ライブラリinstall
sudo apt install ros-${ROS_DISTRO}-plotjuggler-ros
sudo apt install sqlite3


rosbagを作ってみる。

ros2 bag record -o all.bag -a
# or 
ros2 bag record -o all.bag /topic_a /topic_b

# all.bag directory is created
#all.bag
#├── all.bag_0.db3
#└── metadata.yaml

dbファイルと metadata.yamlが作成される。

rosbag再生
$ ros2 bag play all.bag

$ ros2 bag play <rosdir>/ -r 0.2 -s sqlite3 --clock
# -r で倍速指定

# rosbag topic 確認
ros2 bag info <rosbag directry>


・このうちmetadata.yamlはトピック一覧などが書かれた補助情報。
・rosbag本体はall.bag_0.db3。
・この本体がどういった形式になっているかはmetadata.yamlに書かれてる。


sqlite3 で中身を確認

sqlite3 data.bag/data_0.db3
>>>
sqlite> .tables
# messages  metadata  schema    topics
sqlite> .schema


metadata.yamlがない場合

# metadata.yamlを作成
$ ros2 bag reindex <rosbag_dir> -s sqlite3
# >> [INFO] [1737090846.303553233] [rosbag2_cpp]: Reindexing complete.


plotjuggler

PlotJugglerは、データの可視化やリアルタイムモニタリングに特化したros2のツール。簡単に言うと、「データをグラフで見やすく表示し、リアルタイムで観察・分析できるソフトウェア」。



特徴
・リアルタイムモニタリング
データがリアルタイムで変化する様子をその場で観察できるので、ロボットやセンサーの動作確認に便利。
・多くのデータ形式をサポート
CSVファイル、ROS(Robot Operating System)のトピック、JSON形式など、多くのデータ形式を読み込むことができる。


利用シーン
・ロボット開発
センサーや制御システムのデータをリアルタイムで可視化し、動作の改善や問題発見に役立ちます。
・ 機械学習
学習中のデータや予測結果をリアルタイムで確認して、モデルの調整を効率化します。
・実験データの分析
科学や工学の実験結果を手軽に可視化し、分析を行います。


plotjugglerをrosbagで可視化

# download rosbag file from 
gdown -O ~/autoware_map/ 'https://docs.google.com/uc?export=download&id=1sU5wbxlXAfHIksuHjP3PyI2UVED8lZkP'
unzip -d ~/autoware_map/ ~/autoware_map/sample-rosbag.zip
# run plotjuggler 
ros2 run plotjuggler plotjuggler


SWAP realse

# 状況確認
free -h

#現在のSWAP削除
sudo swapoff /swapfile
sudo rm /swapfile

# 新しいswapfile作成
sudo fallocate -l 32G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile

# 反映されてるか確認
free -h

・参考サイト
GitHub - facontidavide/PlotJuggler: The Time Series Visualization Tool that you deserve.

Jetson Orin NanoにROS2(humble)をinstall (2024/10)

Jetson Orin NanoにROS2 (humble)をinstallする方法メモ。

Jetson Orin Nanoのsetup後にROS2をinstallした。
trafalbad.hatenadiary.jp



ROS2には

Humble Hawksbill(LTS)とIron Irwini(latest release)があるが、
安定してるHumble Hawksbill(LTS)をinstallした。

・JetPack 6.1
・ubuntu 22.04
・ROS2 Humble

1. Setup Locale

Locale(言語や単位, 記号, 日付, 通貨などの表示規則の集合のこと)のsetup。

sudo apt update && sudo apt install locales
sudo locale-gen en_US en_US.UTF-8
sudo update-locale LC_ALL=en_US.UTF-8 LANG=en_US.UTF-8
export LANG=en_US.UTF-8


2. ROS2 Repository

ROS2 パッケージ用のRepositoryをinstall

# necessary tools
sudo apt update
sudo apt install -y curl gnupg2 lsb-release software-properties-common

# ROS2 Repository's GPG key
sudo curl -sSL https://raw.githubusercontent.com/ros/rosdistro/master/ros.asc | sudo apt-key add -

# add ROS2 Repository to your Jetson 
sudo sh -c 'echo "deb http://packages.ros.org/ros2/ubuntu $(lsb_release -cs) main" > /etc/apt/sources.list.d/ros2-latest.list'


3. install ROS2 / setup environment

sudo apt update

# install ROS 2 Humble
sudo apt install ros-humble-desktop

# setup env
source /opt/ros/humble/setup.bash
echo "source /opt/ros/humble/setup.bash" >> ~/.bashrc
source ~/.bashrc


Ironのケース

sudo apt install ros-iron-desktop
source /opt/ros/iron/setup.bash
echo "source /opt/ros/iron/setup.bash" >> ~/.bashrc
source ~/.bashrc


4. install tools

ROS 1では rosbuild や catkin が使われてるけど、
ROS2ではbuild toolsにcolconが使われてる。


ROS1ではcatkinを用いたビルドを行っていました.catkinは直接cmakeのみを扱います.一方,ROS2では,colconと呼ばれるメタビルドシステムを用います.colconは依存関係を考慮してパッケージのビルド順を決め,ビルドを実行します.ビルドの方法は各パッケージに任せるので,cmakeによらず複数のビルドタイプを選択可能です(ROS2チュートリアル 体験記 (2/3) | Tokyo Opensource Robotics Kyokai Association)

# install 
sudo apt install python3-colcon-common-extensions

# install resdep
sudo apt install python3-rosdep2
sudo rosdep init
rosdep update

これでROS2のinstallは完了したはず

5 テスト

ROS2を動かしてみる

ros2 run demo_nodes_cpp talker
ros2 run demo_nodes_cpp listener

色々変更はあるけど現時点でわかってるのは

・ROS2のlaunchファイルは,xml形式からpython形式になったということ
・ROS1におけるrospyはrclpyに,roscppはrclcppに変更
・roscoreは必要ない通信形式

6 その他のsetup

Jetson Orin NanoのPerformanceの最大化

sudo nvpmodel -m 0
sudo jetson_clocks

turtlesim

sudo apt update
sudo apt install ros-humble-turtlesim
# Check that the package is installed:
ros2 pkg executables turtlesim
>>
turtlesim draw_square
turtlesim mimic
turtlesim turtle_teleop_key
turtlesim turtlesim_node
# start 
# terminal 1
ros2 run turtlesim turtlesim_node
# terminal 2
ros2 run turtlesim turtle_teleop_key
# terminal 3
rqt_graph


ROS2で試しにbuild

# git clone source code
mkdir -p ~/ros2_ws/src
cd ~/ros2_ws/src
git clone https://github.com/ros2/ros2.git -b humble src/ros2

# install dependencies by rosdep
rosdep install --from-paths src --ignore-src -r -y

# build workspace
colcon build --symlink-install

basicなフォルダ構成はこんな感じ

Jeston@orin:~/Downloads/ros2_ws$ tree
.
├── build
│   └── COLCON_IGNORE
├── install
│   ├── COLCON_IGNORE
│   ├── local_setup.bash
│   ├── local_setup.ps1
│   ├── local_setup.sh
│   ├── _local_setup_util_ps1.py
│   ├── _local_setup_util_sh.py
│   ├── local_setup.zsh
│   ├── setup.bash
│   ├── setup.ps1
│   ├── setup.sh
│   └── setup.zsh
├── log
│   ├── build_2024-10-21_20-32-47
│   │   ├── events.log
│   │   └── logger_all.log
│   ├── COLCON_IGNORE
│   ├── latest -> latest_build
│   └── latest_build -> build_2024-10-21_20-32-47
└── src
    └── ros2
        ├── README.md
        └── ros2.repos

ROS2を無事にinstallできた。

Jetson Orin Nano のセットアップ方法 2024/10

Jetson Orin Nanoでないといろいろ困るので、Orin Nano購入した。なのでsetup方法まとめ


前回の記事のとほとんどからないけど、いくつかパーツが違うので以下を購入した。
trafalbad.hatenadiary.jp


・電源ケーブルDELL/HP用3ピンソケット(メス)⇔2ピンプラグ(オス)

・Amazonベーシック DisplayPort (ディスプレイポート) -- HDMI 変換ケーブル

JetPack


例によってJetPackをJetPack SDKのページからダウンロード。

versionはJetPack==6.1

JetPack 6.1 のデフォルトversion
・ubuntu 22.04.5 (lsb_release -a)
・Python 3.10.12
・opencv 4.8.0 (dpkg -l | grep libopencv)
・cuda 12.6 (nvidia-smi)
・Nvidia driver 540.4.0

GUIの設定

・Google chromiumのインストール
・日本語にキーマボードを変換


CUI で Install & update

# update 
$ sudo apt-get update 
$ sudo apt-get upgrade
$ sudo apt dist-upgrade

# pip3 install
$ sudo apt install gfortran libopenblas-base libopenmpi-dev libopenblas-dev libjpeg-dev libv4l-dev python3-pip
$ sudo pip3 install -U pip
$ sudo pip3 install cython numpy tqdm scikit-learn

# install tree
$ sudo apt install tree


VNC設定

あとで