CPU benchmark - helping the decision

 I was asked to help a friend to chose a laptop (notebook) as his previous one was on the way to die.

REQUIREMENTS

...as first idea were determined to find the best solution on the market matching his needs, such as:

  • strong CPU, at least an Intel Core i7 (at least 9th generation) or an AMD Ryzen7 (3xxx at least),
  • strong GPU, a dedicated one with at least 2 GB video RAM,
  • SSD drive to make it fast, along with DDR4 RAM (size does not matter above 6 GB, as can be increased later),
  • and other wishes which are not relevant from the point of view of this post. 
We all know that to find your favorite device is not an easy task as you plan to have it for several years and hope that it gives you more joy and success than irritation or disappointment.

the CONCEPT

Chosing the right one is in general starts with online search or with real window shopping at first step. The last is for those who wants to feel (touch & watch) the variety of the hardwares, but this was not the case for us.
There are plenty websites with wide spread keyword based search engines. If you are lucky than you know a website which only gathers information from other sites and help customers in finding the list... a vaste list, which at the end is disturbing and increases your doubtfullness whether you make the right choice at the end?
So we have chosen pure and simple comparison of 2 components: CPU and GPU (from which I show the CPU part).
Let's dive in the details of the CPUs.

BENCHMARK

There are plenty of benchmarks on the web to find the 'ultimate comparison' list of the CPUs, or GPUs (or complete laptops), but in which can we trust? How to use and understand benchmark values, Intel's hints may be useful.
He found a table of different benchmark results of some CPUs (an outdated copy of some other website, translated to hungarian, such as this one on NoteBookcheck.net data) and asked me to analyse it. This table contains the following benchmarks:
  • 3DMark, beside other options, it is able to check CPU workload processing capabilities, checks single and multithreading abilities. 
  • Cinebench, which is based on the Cinema 4 Suite and the benchmark have different options, such as 32 (not used here) and 64 bit tests, varied with Single Core and Multi core (hyperthreading) test modes.
  • x264 Pass benchmark determines CPU performance how fast it encodes a 1080p video into the HD x264 video format.
From the viewpoint of the analysis:
  • for all above mentioned benchmark results the higher numbers indicate better performance,
  • single benchmarks can test some properties of the CPU but may not reveal all keypoints,
  • there is a rumour that CPU makers are specialising their CPUs to have the best results with defined test softwares, but for sure they cannot optimise for all. It would be a large benefit for us Users, but it is either impossible or ... results in an expensive hardware. 
So the data analysis have been started through the steps of...

ETL

... jump over this section if you are not interested, see Results***
Well this is not a real ETL (in this post, see for real ETL here***) as the process here will be shown using
  • Microsoft Excel
  • Google Sheets
  • Notepad++
  • Google Data Studio

  1. Data extraction and cleaning



The data of the table was copied to an excel worksheet. But of course everything fall apart as Excel considers that is able to figure out what we want and what we mean by the inserted content (type, the hungarian decimal comma separator, converting numbers to dates,...), but otherwise there would be no need for ETL. :)
I tried other modes, such as Notepad++ but the data on the hungarian site was not well configured, so it turned out even worse. I continued with Excel, saved into csv format.






As you can see data is not well organized and a lot of element were defined in a wrong way.
I switched to NotePad++ to use the Exchange option using regexps.

The cleaning:
  • header removed,
  • separator was selected to be semicolon (so not a strict csv file was made), to eliminate further errors of
  • decimal separator, dot exchanged to comma (safety mode, made row-by-row):  . → , to match my regional settings

detailed ETL steps (involving steps below)

Regrouping and extracting data

  1. The aim was to have numbers only if possible, without physical units. The following changes have been made, removed: 
    • not required units, names, 
    • splitting CPU made, type and model (space and '-' as separators) into 3 columns (first creating empty columns to the right, not to overwrite existing data)
    • splitting CPU model and modifying symbols



    • splitting number of cores and threads
    • splitting turbo min/max frequency values separated by a hyphen (data not used, just for the sake of ETL and real fanatics)
      The result:
    • suspicious number verification and fixing (sometimes by internet search)

  2. Preparation for Machine Learning

    Ultimate conversion of all remaining text based descriptors to numbers.
    The result as data table was not used in the following steps, but have been saved for ML process in a separate csv file. (see in another post)


  3. Data load
    Data was loaded to Excel (new worksheet) and into Google Data Studio as well, for demonstration purposes.

Visualization

None of the below softwares are perfect for the aim, but both can be used to retrieve some information, of course GDS is way better for this analysis purpose. Missing features of the softwares are mentioned in the text.
  1. Excel for the basics

    Depending on the number of cores, the power consumption may vary a lot. This parameter also determines the performance, however the best performance does not go along with the highest power, see that real outlier at 100 W:



    To have faster CPUs, the base and all related communication (BUS) frequency is following an increasing tendency since ... ever. Of course it is not only the frequency, which determines the final efficiency, but it is a highly relevant parameter (as a rule of thumb for any kind of electronical processing and communication). The reached benchmark values (3D Mark in this case) are plotted against maximum turbo frequencies (in MHz), in case the turbo frequency was not defined, the CPU's generally determined frequency was used: 


    The best 3D benchmark result: Intel Core-i9 9980HK and the worst (among those which has data given) Intel Core-i5 4300Y ... of course these are not solid statements and do not count as the aim was something medium amongst the medium level CPUs, for an affordable price. (This is not an Intel ... 9980HK advertisement! Just a statement based on the given data (For fanatics and censorious, I have checked: AMD Ryzen 9 3900X is better now, Sep 2020)!)

    Excel is not appropriate to make charts/plots with labeled data points, which would facilitate finding the best/worst/any item. There are details (data series name, X and Y value), but that is not helpful in all cases and not interactive at all!

    (The word data series in hungarian: "adatsor" has been broken in to two, unknown reason, must be a bug in Excel :) )

    Excel can deal with simple data rows:

    Lots of conceptual plotting errors, useless/incomplete information in the pop-up panel

    It is not what we want, but draws our (data cleaner) attention on missing (zero) values and general things to be aware of.



  2. Data Studio for the fine details
    In this situation data was prepared (structured) with the help of Excel and Notepad++, but of course new Parameters and Fields can be easily defined in Data studio, such as a benchmark results weighted with the power consumption which may be a good feedback on





Blog design

If you wish to move Pages to the top to be present as a Menu.
Ha fent, menü-szerűen szeretnéd megjeleníteni az Oldalakat (Pages).







Save at the bottom, then Save at the bottom right.
Ments alul és Ments még egyszer a jobb alsó sarokban!

To Select Pages to show (note: changed blog color design may uncover hidden pages!):
A megjelenítendő Oldalak (Pages) kiválasztása (angolul):
https://support.google.com/blogger/answer/165955?hl=en

Online course of pure Data Science

timing: March 2020 - June 2020
website: https://data36.com/
useful: Yes! Worth to complete.
pay-service: yes

Linux based webserver (in the cloud, Linode.com) was used to store (SQL server), prepare (ETL processes in bash and Python), analyze (SQL, Python, Google Data Studio, Tableau) data to derive conclusions and to create plots and figures to be involved in presentation files discussed with the tutor of the 6+6 weeks Real life data based DataScience online (video) course. (see details or join the course @ https://data36.com/).
This was a quite robust and well organized course including a large number of lessons, with several datasets to analyze from different aspects. It keeps you every day buzy and excited, which is a benefit that you do not have time to forget what you have just learnt.
Here are the major tools used while the datasets were prepared, analyzed and then discussed:



Large datasets (10000+ data rows in several subsets) were used to make User statistics (e.g. churn) and business analysis (costs, income, cashflow, funnel analysis) of a webpage, or an online app. Some of the results of my work can be checked as a presentation file (Google slides, shared).

See further details in the file. If you are still missing something, please get in contact with me or with the tutor! I highly recommend him and his course (without any benefit for me).

Online Data Science kurzus

időtartam: 2020 március - június
webhely: https://data36.com/
oklevél: igen
hasznosság: Igen! Nagyon megéri.
fizetős: igen

OKLEVÉL

Linux alapú webszervert (felhőben, Linode.com) használtam adat-tárolásra (SQL szerver), előkészítésre (ETL folyamatok: bash és Python), elemzésre (SQL, Python, Google Data Studio) következtetések levonására, valamint diagramok és ábrák készítésére (rendszerkörnyezet telepítési lista). Nagyszerű volt részt venni ebben a 6 + 6 hetes, valós adatokon alapuló DataScience online (video) tanfolyamban, ami során készült prezentációs fájlokat bemutattuk a kurzus oktatójának. (Részletek, vagy a tanfolyamhoz csatlakozás miatt keresd fel a https://data36.com/ oldalt).

Ez egy meglehetősen összetett tanfolyam volt, amely sokféle tananyagból állt, több elemzendő adatkészlettel, különböző vizsgálati szempontokkal. Szinte minden nap adott elfoglaltságot, ami igazából előny, mert így nem felejted el könnyen az épp megtanult dolgokat.

A főbb eszközök, amelyeket az adatkészletek feldolgozása (előkészítés, tisztítás, ...), elemzése, majd megbeszélése során használtam:



A legutolsó feladatban egy üzleti elemzést (business analysis) kellett végezni (10 000+ adatból álló) több részletben megadott adatbázison, ami egy kitalált weboldal működését tükrözte. Ennek az elemzésnek a részleteit és eredményeit megtaláljátok itt egy prezentációs fájlban.

További részletek a fájlban találhatók. Ha úgy érzed hiányzik valamilyen információ, kérlek, lépj kapcsolatba velem vagy az oktatóval! Nagyon ajánlom őt és a tanfolyamát (mindenféle ellenszolgáltatás nélküli az ajánlás)!

JDS Academy - course material #2

JDS Academy - course material #2


Supplementary material: Cheat Sheets

  • SQL,
  • Python,
  • Bash



Additional JDS course material:
table of content of A/B testing.

WEEK 0

Installation of a Linux webserver  on Linode cloud + I made a Linux workstation, both with Ubuntu.

WEEK 1

# Module 1: Introduction to Data Science

  • What is Data Science? Case study.
  • Clarification of AI, ML, big data, deep learning concepts.

# Module 2: How to become a Data Scientist

  • Soft skills, mindset, time commitment, roadmap, learning curve.

# Module 3: SQL Introduction - European accidents dataset

  • Introduction to SQL Workbench and pgAdmin, installation.
  • SQL server configuration
  • Importing data into SQL, data analysis.

# Module 4: SQL Basics + Simple Queries

  • Basic SQL exercises.

# Module 5: SQL WHERE with Multiple Filter Conditions - European accidents dataset

  • Advanced filtering techniques.

# Module 6: SQL WHERE + ORDER BY - European accidents dataset

  • Combination of sorting and filtering.

# Module 7: SQL Functions (COUNT, SUM, AVG, MIN, MAX) - European accidents dataset

  • Basic Aggregation Functions.

WEEK 2

# Module 1: SQL Functions + GROUP BY

  • Grouping and Aggregation.

# Module 2: SQL Table Joining (JOIN)

  • Table Joining Techniques.

# Module 3: SQL Subqueries + HAVING

  • Nested Queries and Conditional Filtering.

# Module 4: Extra Task (Case Study) - Mobile App User and Business Data Analysis

  • Data-Driven Tasks.

# Module 5: Data Analysis Methodologies

# Vault: SQL Exercises

  • Interview-Preparing SQL Tasks.
    • A/B test data, 
    • Solar panel factory production data analysis, 
    • Travel blog User and business analysis.

WEEK 3

# 0. Module: Setting up your own server

  • Installing and configuring a database server.
    Note: I have done on the 0th week. Linode webserver + Linux workstation, Ubuntu

# 1. Module: Bash Basics

  • Basics of Bash commands and scripts. ETL with bash

# 2. Module: Bash Basics continued

  • Additional practical tasks. ETL with bash

# 3. Module: Scripts and automation in Bash

  • Automation techniques.

# 4. Module: Data collection

WEEK 4

# 1. Module: Python Introduction, Variables, Data structures

  • Jupyter Notebook management, variables and structures.

# 2. Module: Python Functions, If Branches, For Loops

  • Basic Python programming techniques.

# Module 3: Python Practice

  • Various python practice tasks, basic operations/logical functions.

# Module 4: Statistics

WEEK 5

# Module 1: Python + Analytics: Pandas basics

  • Data Management with Pandas.

# Module 2: Pandas GroupBy, Functions, Sorting

  • Advanced data management techniques.

# Module 3: Data Visualization with Python + Pandas

  • Creating graphs and visualizations.

# Module 4: Predictive Analytics

WEEK 6

# Module 1: Machine Learning Examples in Python

  • Linear and Polynomial Regression, 
  • Random Forest, 
  • Deep Learning.

# Module 2: Data Presentation


JDS Académia 2

A Junior Data Scientist Akadémia (JDS) kurzus anyaga

 

 Cheat Sheet-ek:

  • SQL,
  • Python,
  • Bash


 1. HÉT

# 1. Modul: Bevezetés a Data Science világába

- Mi az a Data Science? Esettanulmány.

- AI, ML, big data, deep learning fogalmak tisztázása.

# 2. Modul: How to Become a Data Scientist

- Soft skills, mindset, időráfordítás, roadmap, tanulási görbe.

# 3. Modul: SQL Bevezetés - European accidents dataset

- SQL Workbench és pgAdmin bemutatása, telepítés.

- SQL szerver konfigurálás

- Adatok importálása SQL-be, adatelemzés.

# 4. Modul: SQL Alapok + Egyszerű Lekérdezések

- Alapvető SQL gyakorlatok.

# 5. Modul: SQL WHERE Több Szűrőfeltétellel - European accidents dataset

- Haladó szűrési technikák.

# 6. Modul: SQL WHERE + ORDER BY - European accidents dataset

- Rendezés és szűrés kombinációja.

# 7. Modul: SQL Függvények (COUNT, SUM, AVG, MIN, MAX)

- Alapvető aggregáló függvények.


 2. HÉT

# 1. Modul: SQL Függvények + GROUP BY

- Csoportosítás és aggregálás.

# 2. Modul: SQL Táblák Összekapcsolása (JOIN)

- Táblák összekapcsolási technikái.

# 3. Modul: SQL Subquery-k + HAVING

- Beágyazott lekérdezések és feltételes szűrés.

# 4. Modul: Extra Feladat (Esettanulmány) - Mobil app User és üzleti adatelemzés

- Adatvezérelt feladatok.

# 5. Modul: Adatelemző módszertanok

# VAULT: SQL Gyakorlatok

- Interjúra készítő SQL feladatok.

A/B test adatok, Napelem gyár gyártási adatelemzés, Utazó blog User és üzleti elemzés


 3. HÉT

# 0. Modul: Saját Szerver Beállítása

- Adatbázis-szerver telepítése és konfigurálása.

# 1. Modul: Bash Alapok

- Bash parancsok és szkriptek alapjai. ETL bash-sel

# 2. Modul: Bash Alapok Folytatás

- További gyakorlati feladatok. ETL bash-sel

# 3. Modul: Script-ek és automatizálás Bash-ben

- Automatizálási technikák.

# 4. Modul: Adatgyűjtés


 4. HÉT

# 1. Modul: Python Bevezetés, Változók, Adatszerkezetek

- Jupyter Notebook kezelése, változók és struktúrák.

# 2. Modul: Python Függvények, If elágazások, For Loop-ok

- Alapvető Python programozási technikák.

# 3. Modul: Python Gyakorlás

- Különböző python gyakorlófeladatok, alap műveleti/logikai függvények.

# 4. Modul: Statisztika


 5. HÉT

# 1. Modul: Python + Analitika: Pandas Alapok

- Adatkezelés Pandas-szal.

# 2. Modul: Pandas GroupBy, Függvények, Sorting

- Haladó adatkezelési technikák.

# 3. Modul: Adatvizualizáció Python-nal + Pandas-szal

- Grafikonok és vizualizációk készítése.

# 4. Modul: Prediktív Analitika


 6. HÉT

# 1. Modul: Machine Learning Példák Python-ban

- Lineáris és Polinomiális Regresszió, Random Forest, Deep Learning.

# 2. Modul: Adatok Prezentálása


 BÓNUSZ

- A/B tesztelős kurzus hozzáférés.

Snowflake universe, part #6 - Forecasting2

Forecasting with built-in ML module Further posts in  Snowflake  topic SnowFlake universe, part#1 SnowFlake, part#2 SnowPark Notebook Snow...