🐧 Linux TippsDebian 11 Long Term Support reaches end-of-life(31.08.2026 um 02:00 Uhr)
🐧 Linux TippsUpdated Debian 13: 13.7 released(12.09.2026 um 02:00 Uhr)
🕵️ SicherheitslückenUSN-8741-1: Flatpak vulnerabilities(10.09.2026 um 10:44 Uhr)
🕵️ SicherheitslückenUSN-8742-1: Netty vulnerability(10.09.2026 um 11:01 Uhr)
🕵️ SicherheitslückenUSN-8737-2: GNU C Library vulnerabilities(10.09.2026 um 13:25 Uhr)
🕵️ SicherheitslückenUSN-8743-1: PHP vulnerabilities(10.09.2026 um 13:48 Uhr)
🕵️ SicherheitslückenUSN-8744-1: Python vulnerabilities(10.09.2026 um 15:53 Uhr)
🐧 Linux TippsUSN-8748-1: Linux kernel (NVIDIA) vulnerabilities(10.09.2026 um 17:32 Uhr)
🕵️ SicherheitslückenUSN-8745-1: KissFFT vulnerabilities(10.09.2026 um 17:36 Uhr)
🕵️ SicherheitslückenUSN-8746-1: libEBML vulnerability(10.09.2026 um 17:48 Uhr)
🐧 Linux TippsDebian 11 Long Term Support reaches end-of-life(31.08.2026 um 02:00 Uhr)
🐧 Linux TippsUpdated Debian 13: 13.7 released(12.09.2026 um 02:00 Uhr)
🕵️ SicherheitslückenUSN-8741-1: Flatpak vulnerabilities(10.09.2026 um 10:44 Uhr)
🕵️ SicherheitslückenUSN-8742-1: Netty vulnerability(10.09.2026 um 11:01 Uhr)
🕵️ SicherheitslückenUSN-8737-2: GNU C Library vulnerabilities(10.09.2026 um 13:25 Uhr)
🕵️ SicherheitslückenUSN-8743-1: PHP vulnerabilities(10.09.2026 um 13:48 Uhr)
🕵️ SicherheitslückenUSN-8744-1: Python vulnerabilities(10.09.2026 um 15:53 Uhr)
🐧 Linux TippsUSN-8748-1: Linux kernel (NVIDIA) vulnerabilities(10.09.2026 um 17:32 Uhr)
🕵️ SicherheitslückenUSN-8745-1: KissFFT vulnerabilities(10.09.2026 um 17:36 Uhr)
🕵️ SicherheitslückenUSN-8746-1: libEBML vulnerability(10.09.2026 um 17:48 Uhr)

🔧 Programmierung 🕛 vor 1 Jahr 5 Min Lesezeit
0

Study Notes 3.3.1: BigQuery Machine Learning (BQML)

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht
📺
dev.to







  1. Data Pipeline:


    • Data collection and ingestion

    • Data processing and feature engineering

    • Dataset splitting (training/test)

    • Model building

    • Hyperparameter optimization

    • Model validation

    • Model deployment



  2. Feature Processing Options:


    • Automatic preprocessing:


      • Standardization of numeric fields

      • One-hot encoding of categorical fields

      • Multi-hot encoding of arrays



    • Manual preprocessing:


      • Bucketization

      • Polynomial expansion

      • Feature cross

      • N-grams

      • Min-max scaling










Algorithm Selection Guide



Algorithm Selection Guide



Based on use case:




  1. Value Prediction (e.g., sales figures, stock prices):


    • Linear regression

    • Boosted tree

    • AutoML

    • DNN regressor

    • Wide and deep regression



  2. Customer Segmentation:


    • K-means clustering



  3. Classification (e.g., spam detection):


    • Logistic regression

    • Boosted tree classifiers

    • AutoML tables








Model Creation and Evaluation




  1. Model Creation Steps:


    • Define model type and algorithm

    • Specify input label columns

    • Choose data split method (e.g., auto split)

    • Configure any necessary hyperparameters



  2. Evaluation Capabilities:


    • Built-in evaluation metrics (e.g., mean squared error, mean absolute error)

    • Feature importance analysis

    • Model explanation tools

    • Prediction capabilities with ML.PREDICT

    • Explain and predict functionality for feature importance








Hyperparameter Tuning Options



Available parameters for optimization:




  • Number of trials

  • Max parallel trials

  • L1 regression hyperparameter

  • L2 regression hyperparameter

  • Learning rate strategy

  • Early stop settings

  • Custom learning rates






Best Practices




  1. Data Preparation:


    • Proper type casting for categorical variables

    • Appropriate feature selection

    • Data cleaning and filtering

    • Consider automatic vs manual preprocessing needs



  2. Model Optimization:


    • Use evaluation metrics to guide improvements

    • Leverage hyperparameter tuning options

    • Consider multiple algorithm options for best results

    • Monitor model performance and costs





The focus is on using SQL and leveraging BigQuery's built-in capabilities, minimizing the need for external tools or programming languages like Python.



I. Introduction to BQML:




  • BQML is designed for data analysts and managers familiar with SQL and basic machine learning concepts.

  • It allows model building directly within the data warehouse, eliminating the need for data export.


  • Pricing: BigQuery offers a free tier for data storage, queries, and model creation. Beyond the free tier, costs are incurred per terabyte of data processed, varying by model type (e.g., linear regression, boosted trees, AutoML). Regional pricing may vary.



II. Machine Learning Development Steps & BigQuery's Role:



The typical ML development process involves:





  1. Data Collection: Data is assumed to be already in BigQuery.


  2. Data Processing/Feature Engineering: BigQuery supports both automatic and manual feature preprocessing.



    • Automatic Preprocessing: Includes standardization of numeric fields, one-hot encoding of categorical fields, and multi-hot encoding of arrays.


    • Manual Preprocessing: Provides options like bucketization, polynomial expansion, feature cross, n-grams, min-max scaling, etc.




  3. Model Building: Choosing the algorithm and tuning hyperparameters. BigQuery offers various algorithms and hyperparameter tuning capabilities.


  4. Model Validation: Evaluating model performance using various metrics. BigQuery provides many error matrices for validation.


  5. Model Deployment: BigQuery allows model deployment using Docker images.



III. Algorithm Selection:



BigQuery provides a range of algorithms categorized by use case:





  • Predicting Values (Regression): Linear Regression, Boosted Trees, AutoML DNN Regressor, Wide and Deep Regression.


  • Customer Segmentation: K-Means.


  • Classification (Predicting Categories): Logistic Regression, Boosted Trees Classifier, AutoML Tables.



IV. Building a Linear Regression Model (Example):



The video demonstrates building a linear regression model to predict tip amount using the Yellow Trip Data dataset.





  1. Data Selection & Preparation: A SQL query selects relevant features (passenger count, trip distance, pickup/drop-off location IDs, payment type, fare amount, tolls). Crucially, categorical features (location IDs, payment type) are cast to STRING type. This allows BigQuery's automatic preprocessing to perform one-hot encoding. A new table is created with these modified data types. Filtering is also applied (e.g., fare_amount not equal to zero).


  2. Model Creation: The CREATE MODEL statement is used:



    • model_name: A name for the model (e.g., tip_model).


    • model_type: LINEAR_REGRESSION.


    • input_label_cols: The target variable (e.g., tip_amount).


    • data_split_method: AUTO_SPLIT (for training and evaluation data).




  3. Model Inspection: ML.GET_MODEL displays model details (type, training/evaluation data, loss, duration, evaluation metrics).


  4. Feature Information: ML.FEATURE_INFO shows how features were processed (numeric vs. categorical, min/max, mean, category counts).


  5. Model Evaluation: ML.EVALUATE assesses model performance on the evaluation data (using metrics like mean squared error, mean absolute error).


  6. Prediction: ML.PREDICT generates predictions on a dataset, adding a predicted_tip_amount column.


  7. Explain & Predict: ML.EXPLAIN_PREDICT identifies the most influential features for predictions.


  8. Hyperparameter Tuning: BigQuery offers various hyperparameters for tuning different models (e.g., number of trials, L1/L2 regularization for linear regression). Refer to the documentation for a full list.



V. Key Takeaways:




  • BQML simplifies machine learning by using SQL and integrating model building into the data warehouse.

  • Automatic feature preprocessing handles common transformations like one-hot encoding.

  • BigQuery provides a range of algorithms and tools for model evaluation and hyperparameter tuning.

  • The example demonstrates a basic linear regression model; more complex models and tuning are possible.

  • Understanding data types and how they are handled by BigQuery's preprocessing is crucial.



VI. Further Exploration:




  • Explore the full range of algorithms and hyperparameter tuning options available in BigQuery.

  • Investigate manual feature engineering techniques to improve model performance.

  • Learn about model deployment options using Docker images.

  • Consider the cost implications of using BQML beyond the free tier.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Debian 11 Long Term Support reaches end-of-life
1 Quelle
Updated Debian 13: 13.7 released
1 Quelle
USN-8741-1: Flatpak vulnerabilities
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Study Notes 3.3.1: BigQuery Machine Learning (BQML)

Thematisch verwandte Begriffe: Study, Notes, BigQuery, Machine · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...