Least squares optimization in SciPy - Time & Space Complexity
Start learning this pattern below
Jump into concepts and practice - no test required
We want to understand how the time needed to solve a least squares problem grows as the data size increases.
How does the number of calculations change when we have more data points or variables?
Analyze the time complexity of the following code snippet.
import numpy as np
from scipy.optimize import least_squares
def model(x, t):
return x[0] * np.exp(-x[1] * t)
t = np.linspace(0, 10, 100)
y = model([2.5, 1.3], t) + 0.1 * np.random.randn(100)
res = least_squares(lambda x: model(x, t) - y, x0=[1, 1])
This code fits a model to data by minimizing the difference between the model and observed points using least squares.
Identify the loops, recursion, array traversals that repeat.
- Primary operation: Calculating residuals (differences) for all data points in each iteration.
- How many times: For each iteration, the residuals are computed over all data points; iterations repeat until convergence.
As the number of data points grows, the calculations for residuals increase proportionally each iteration.
| Input Size (n) | Approx. Operations |
|---|---|
| 10 | About 10 residual calculations per iteration |
| 100 | About 100 residual calculations per iteration |
| 1000 | About 1000 residual calculations per iteration |
Pattern observation: The work grows roughly in direct proportion to the number of data points.
Time Complexity: O(n * k)
This means the time grows linearly with the number of data points (n) and the number of iterations (k) needed to find the best fit.
[X] Wrong: "The time depends only on the number of data points, not on the number of iterations."
[OK] Correct: Each iteration requires recalculating residuals for all points, so more iterations multiply the total work.
Understanding how least squares optimization scales helps you explain performance when fitting models to data, a common task in data science and machine learning.
"What if the model had more parameters to estimate? How would the time complexity change?"
Practice
scipy.optimize.least_squares in data fitting?Solution
Step 1: Understand the purpose of least squares
Least squares optimization aims to find parameters that reduce the error between predicted and actual data.Step 2: Connect to
This function specifically minimizes the sum of squared residuals, which are differences between model and data.scipy.optimize.least_squaresFinal Answer:
To find parameters that minimize the difference between the model and data -> Option CQuick Check:
Least squares = minimize difference [OK]
- Thinking it maximizes difference
- Confusing with sorting or random selection
- Assuming it changes data order
scipy.optimize.least_squares with a residual function fun and initial guess x0?Solution
Step 1: Check the function signature
The correct call isleast_squares(fun, x0)wherefunis the residual function andx0is the initial guess.Step 2: Verify argument order
Arguments must be in order: first the function, then the initial guess.Final Answer:
least_squares(fun, x0) -> Option DQuick Check:
Function first, initial guess second [OK]
- Swapping argument order
- Using keyword arguments incorrectly
- Omitting the initial guess
import numpy as np
from scipy.optimize import least_squares
def residuals(x):
return np.array([2*x[0] - 4, x[1] + 3])
result = least_squares(residuals, [0, 0])
print(result.x)Solution
Step 1: Solve residual equations for zero residuals
Set residuals to zero: 2*x0 - 4 = 0 => x0 = 2; x1 + 3 = 0 => x1 = -3.Step 2: Confirm least_squares finds these values
The optimizer finds x = [2, -3] minimizing residuals to zero.Final Answer:
[2.0, -3.0] -> Option BQuick Check:
2*2-4=0 and -3+3=0 [OK]
- Not solving equations correctly
- Confusing signs in residuals
- Assuming initial guess is output
least_squares:from scipy.optimize import least_squares
def fun(x):
return x**2 - 4
result = least_squares(fun)
print(result.x)Solution
Step 1: Check least_squares function call
The call lacks the required initial guess argumentx0.Step 2: Confirm residual function and print are correct
The residual function returns an array-like (scalar is acceptable as 1D array), and print syntax is valid.Final Answer:
Missing initial guess argument in least_squares call -> Option AQuick Check:
least_squares needs initial guess [OK]
- Forgetting initial guess
- Thinking scalar residuals cause error
- Misreading print syntax
y = mx + c to data points x = [1, 2, 3] and y = [2, 3, 5] using least_squares. Which residual function correctly represents the difference between observed and predicted values?Solution
Step 1: Understand residual definition
Residuals are predicted minus observed values: (model - data).Step 2: Check each function
def residuals(p):\n m, c = p\n return [(m*x[i] + c) - y[i] for i in range(len(x))] returns (m*x + c) - y, matching predicted minus observed.Final Answer:
def residuals(p):\n m, c = p\n return [(m*x[i] + c) - y[i] for i in range(len(x))] -> Option AQuick Check:
Residual = predicted - observed [OK]
- Swapping predicted and observed in residuals
- Adding instead of subtracting values
- Incorrect sign on intercept
